The 9 box grid is a 3x3 chart that plots employees on two axes: how well they are performing now, and how far they could go. Nine cells, nine different conversations. HR teams use it to decide who to promote, who to develop, and who to move.
It has been standard equipment in talent reviews for decades, and it attracts steady criticism the whole time. Both of those hold for the same reason. The grid asks managers to rate an employee’s potential, and potential is much harder to judge than performance.
This guide covers what each of the nine boxes means, how to run the review, where the model actually came from, and what the research says about its weakest part.
What is the 9 box grid?
The 9 box grid is a talent assessment tool that rates each employee on performance and potential, usually on a low / medium / high scale, then places them in the cell where the two ratings meet. Three levels on each axis produce nine boxes. You will also see it called the 9 box matrix, the nine box matrix, the 9 box model or the 9 box talent review.
The output is a single picture of your workforce. Instead of nine separate manager opinions expressed in different language, you get one grid where every name sits somewhere, and where the disagreements between managers are visible.
Most companies use it for three things:
- Succession planning — identifying who could step into a critical role, and how soon.
- Development planning — deciding where coaching, training and stretch assignments go.
- Calibration — seeing how performance and potential are distributed across teams, and where the ratings look implausible.
It is a discussion tool, not a measurement instrument. That distinction is the source of most of the trouble people have with it.
Where the 9 box grid came from, and why "McKinsey 9 box" means two different things
Search for the McKinsey 9 box talent matrix and you will find dozens of pages crediting McKinsey with inventing a talent tool in the 1970s. That is not quite what happened, and the difference matters.
By McKinsey’s own account, the GE–McKinsey nine-box framework was built in the early 1970s to help General Electric decide where to put its money across more than 150 business units. The two axes were the attractiveness of the industry and the business unit’s competitive strength within it. It rated businesses, not people.
The talent grid borrowed the 3x3 shape and swapped the axes for performance and potential. When someone searches for the 9 box model McKinsey, they are almost always looking for the talent version, which is a later adaptation of a portfolio planning tool. Who first made that adaptation, and when, is not well documented, so treat any confident date you read with suspicion.
This is more than a footnote. The original logic was about capital allocation: back the units that will win, hold the steady ones, get out of the rest. Applied to a grid of people, that same logic says invest in some employees and divest others. Much of the discomfort managers feel with the bottom row of the grid comes from the fact that the framework was never designed for human beings in the first place.
The two axes: performance and potential
Performance means results already delivered. Goals hit, work shipped, quality maintained, behaviour demonstrated. It is backward-looking and there is evidence for it: review ratings, delivery records, customer outcomes.
Potential means the capacity to succeed in a bigger or different role. It is forward-looking, which means it is a forecast. By definition there is no evidence yet, because the person has not done the job.
The grid plots a measurement against a prediction and draws them as if they carry equal weight. They do not. Keep that asymmetry in mind through the rest of this guide, because the research at the end of it is entirely about the second axis.
One practical note: templates disagree about which axis is which. Some put performance on the horizontal axis, some on the vertical. Label your axes explicitly before you circulate a grid, or half the room will read it upside down.
The nine boxes and what each one means
Box names vary between companies. What follows describes what each placement usually indicates and the action that normally follows. Treat the names as shorthand, not as job titles.

High performance, high potential
Your strongest people, delivering now and capable of more. This is the succession pool. They need stretch assignments, exposure to senior decisions and a clear line of sight to the next role. They are also your highest flight risk, because they have options.
High performance, medium potential
Reliable strong performers who could grow one level, but not indefinitely. Promote them into that next level when it opens, and invest in deepening their expertise rather than broadening it. Often the most under-appreciated group on the grid.
High performance, low potential
Experts and long-serving specialists who are excellent at the job they hold and do not want or suit a larger one. They carry institutional knowledge and keep operations running. Pay them properly, give them technical or mentoring tracks, and stop treating a flat career path as a failure.
Medium performance, high potential
People with visible capability who have not yet produced the results to match. Usually the cause is fixable: a role that does not fit, an absent manager, unclear goals or too little time in seat. Diagnose the cause before you invest in coaching.
Medium performance, medium potential
The middle of the grid, and normally the largest group. Solid contributors meeting expectations. They need clear goals and regular feedback rather than a development programme. Resist the urge to write them off as average; most of your output comes from this box.
Medium performance, low potential
Steady in the current role with limited appetite for more. Keep expectations explicit, address specific gaps as they appear, and do not spend development budget here by default. Check whether the low potential rating reflects the person or a manager who has stopped looking.
Low performance, high potential
The box that most often signals a management problem rather than an employee problem. Strong raw capability with weak results usually points to role misfit, a health or personal issue, or a manager who has not set direction. Have the conversation before you have the performance review.
Low performance, medium potential
Underperforming with some room to grow. Decide explicitly between a performance improvement plan with a defined timeline, a move to a role that fits better, or an exit. Leaving someone in this box for another year is the default and the worst option.
Low performance, low potential
Results are not there and there is no visible path to them. This needs a documented, honest conversation quickly. Before you act, confirm the rating is not the product of one manager’s view alone, because this is the box where a single biased judgement does the most damage.
How to run a 9 box talent review
The grid is the easy part. The process around it decides whether the output is worth anything.
Agree what performance and potential mean before you rate anyone
Write down the definitions and the scale. What counts as high performance at each level? What specific, observable behaviours count as high potential? If you skip this step, managers apply their own private definitions and the grid becomes a map of manager personalities.
Collect the evidence first
Pull goal attainment, review ratings, manager notes and whatever your 360 feedback software has already captured. Do it before the meeting, not during it. Every rating should have something behind it that someone else could check.
Run a calibration session
Get managers in one room and compare ratings across teams. This is where you find the manager who rates everyone high and the one who rates nobody high. Performance calibration is the step that turns nine private opinions into one usable grid, and it is the step most often cut for time.
Place people, then write down why
Record a one-line reason for every placement, especially on the potential axis. If a placement cannot be explained in a sentence with evidence in it, it is a guess. Those notes are also what let you audit the grid later.
Turn placements into actions with owners
A box with no action attached is just a label. Every name should leave the meeting with a next step, a person responsible and a date. Link each one to an individual development plan so the employee sees development rather than classification.
Set an expiry date
A placement describes a moment. People change roles, managers and circumstances. Re-run the grid on a schedule and treat last cycle’s boxes as history, not as a starting position.
{{banner-7="/banner-page"}}
What the research says about rating potential
This is the part most guides to the 9 box grid leave out, and it is the part that should change how you run the review.
Alan Benson, Danielle Li and Kelly Shue studied performance and potential ratings for roughly 30,000 management-track employees at a large North American retail chain between 2009 and 2015. Their paper, “Potential” and the Gender Promotion Gap, was published in the American Economic Review. Three findings bear directly on the grid.
Potential ratings drove promotion far harder than performance ratings. Moving from a medium to a high potential rating raised an employee’s likelihood of promotion by 75 percent. An equivalent improvement in the performance rating raised it by 27 percent. The forecast axis drove promotion decisions roughly three times as hard as the axis with evidence behind it.
The potential axis carried a measurable gender gap. Women received potential ratings 8.3 percent lower than men, while earning higher performance ratings than men. The two axes pointed in opposite directions for the same people.
The gap in potential ratings explained much of the gap in promotions. Differences in potential ratings explained up to half of the overall gender promotion gap at the company. Women given low potential ratings went on to outperform those ratings, and were still rated lower on potential the following year.
One caveat worth stating: this is one large employer in one sector over six years, not a finding about every company. But the mechanism it describes is the mechanism the 9 box grid runs on. If the vertical axis carries most of the decision weight while resting on the least evidence, that is a design problem in any organisation using it.
Is the 9 box grid outdated?
The 3x3 shape is not the problem. Three specific things are.
- The potential rating is a forecast dressed as an assessment. Managers are asked to predict performance in a job the person has not held, usually without a definition of what they are predicting.
- Labels persist. Once someone has been described as low potential in a room full of senior managers, that description follows them into the next cycle regardless of what they do.
- A snapshot becomes a sentence. An annual placement describes one moment and then gets used as a standing verdict for twelve months.
What the grid is still genuinely good for: forcing a structured conversation that would otherwise not happen, exposing where managers apply different standards, and showing the distribution of talent across a company rather than one team at a time. Those are real benefits and most alternatives do not deliver them.
What it should not be used for: deciding a promotion on its own, ranking people against a forced distribution, or predicting any individual’s future with confidence.
So the honest answer is that the grid is a useful meeting structure being widely misused as a measurement system. Fixing how you use it is a smaller job than replacing it.
How to make the 9 box grid more reliable
Five changes, in the order they are worth making.
- Replace “potential” with “readiness for a named role”. Instead of asking how much potential someone has, ask whether they are ready for a specific named role within a stated timeframe: ready now, ready in one to two years, or not on this path. A narrower question produces an answer you can check.
- Make evidence mandatory on the potential axis. Require a concrete example for every potential rating above medium. No example, no rating. This single rule removes most of the gut-feel placements.
- Rate the two axes in separate passes. Rate every employee on performance first, close that discussion, then start on potential. Rating both at once lets strong performance quietly become evidence of potential, which is the error the research above describes.
- Audit the distribution before you sign it off. Before finalising, look at how potential ratings are distributed by gender, tenure, team and manager. If one group clusters low on potential while holding up on performance, you have found a rating problem, not a talent problem.
- Give placements an expiry date. Set a review date on every placement and delete last cycle’s grid from the conversation when the new one opens.
Whether to tell an employee which box they are in is a real decision with arguments on both sides, and it depends on how much your managers trust each other’s ratings. What matters either way is that the employee leaves with a specific development action rather than a category.
How to build a 9 box grid in Excel or Google Sheets
You do not need software to run your first cycle. A spreadsheet is enough, and building it yourself forces the definitions to be explicit.
Set up one row per employee with these columns:
- Name, role, level and manager.
- Performance rating, 1 to 3.
- Potential or readiness rating, 1 to 3.
- Evidence — one line per axis, written before the calibration meeting.
- Box — derived, not typed. Join the two ratings so the cell fills itself and nobody can quietly override it.
- Action, owner and review date.
To see it as a grid rather than a list, chart the two rating columns as a scatter plot, set both axes to run from 0.5 to 3.5, and add gridlines at 1.5 and 2.5. That produces nine visual zones with your people plotted inside them. Add a little random jitter to each point if names overlap.
Two habits that make the sheet worth keeping: put the evidence in the sheet itself rather than a separate document, and copy the whole tab at the start of each cycle so you can compare placements over time. A grid you can only see once tells you far less than three cycles side by side.
If you want the succession angle specifically, the succession planning template covers how the grid feeds a named successor list.
Where the 9 box grid fits with your other tools
The grid consumes data rather than producing it. It is only as good as the performance information feeding it, which is why it tends to fail in companies without a working review cycle.
In practice it sits between three other things: a review process that generates the performance ratings, a talent management process that turns placements into development, and HR analytics software that shows you the distribution across the company. Get the first one working before you run a grid at all.
If you are choosing a tool rather than a spreadsheet, performance appraisal software increasingly includes the grid as a view over review data instead of a separate exercise. Effy AI does this: reviews produce the performance and potential scores, and heatmaps, score trends and a 9 box view of performance against potential are generated from them, with AI-assisted calibration flagging rating drift between managers and teams. It is free for up to five users, and 360 reviews start at $3 per seat per month billed annually.
For a different lens on the same question, the skill will matrix plots capability against motivation instead of performance against potential, and is often more useful for individual coaching conversations than for company-wide planning.
Using the grid without over-trusting it
The 9 box grid earns its place as a way to get managers into one room, using one vocabulary, making their disagreements visible. That is genuinely hard to achieve any other way, and it is why the tool has outlived so many attempts to retire it.
What it cannot do is predict individual futures. The evidence says the potential axis carries most of the decision weight and the least justification, and that it goes wrong in patterned rather than random ways. Narrow the question to readiness, demand evidence, check the distribution, and let placements expire. Do that and the grid becomes a useful structure. Skip it and you have built a machine for turning manager instinct into career outcomes.
