Microsoft Project Online retires September 30, 2026, migrate to a modern platform before it's too late.Start migration
Back to BlogStory Points and Velocity: What They Measure and What They Don't
Fundamentals

Story Points and Velocity: What They Measure and What They Don't

Story points and velocity get treated as productivity scores. They aren't. Here's what each actually measures, and why turning either into a KPI backfires.

Onplana TeamJuly 19, 20269 min read

Here's the version of this conversation that happens in almost every retro. Velocity dropped from 38 to 31 last sprint, and someone on the leadership call wants to know why the team slowed down. The honest answer, that two people were out sick and a story turned out to depend on an API that wasn't ready, gets waved off. What the room actually wants is a number that goes up every sprint, forever. That number doesn't exist, and chasing it is what breaks agile teams that were otherwise doing fine.

Story points and velocity are two of the most misused numbers in project management, not because the concepts are complicated, but because they get treated as if they measure something they were never built to measure. Story points estimate relative effort. Velocity is a planning tool for one team's own forecasting. Neither is a productivity score, and the moment either gets treated as one, the numbers stop meaning anything.

TL;DR. Story points measure relative size and complexity compared to other work the team has already estimated, not hours. Velocity is the average points a specific team completes per sprint, useful only for that team's own capacity planning. Turning velocity into a target makes teams inflate points instead of ship more work, a straightforward case of Goodhart's Law. Comparing velocity across teams is meaningless because every team's point scale is calibrated independently. The fix is estimating relative to a shared reference story, protecting the number from becoming a KPI, and reading trend direction instead of chasing an absolute target.

What Story Points and Velocity Actually Measure

A story point is a relative unit. It answers "how big is this compared to work we've already sized," not "how many hours will this take." A team picks a small, well-understood story as a reference, typically worth 1 to 3 points, and every subsequent estimate gets compared against it: is the new story about the same size, twice as big, a quarter as big. That relative comparison is the entire mechanism. It works because humans are reliably bad at estimating absolute duration but reasonably good at estimating relative size, the same reason it's easier to say "this bag is heavier than that one" than to guess either bag's weight in kilograms.

Velocity is the downstream number: the total story points a team completes in a sprint, typically averaged over the last three to five sprints to smooth out normal variance. According to the Agile Alliance's definition of velocity, at the end of each iteration the team adds up the effort estimates for stories completed during that iteration, and that total is the velocity, used to forecast how much remaining work the team can realistically absorb in future sprints based on past throughput. Both numbers exist to answer one question: how much work can this specific team reliably plan into a sprint. Neither was designed to answer "is this team working hard enough," and neither holds up when it gets asked to.

Why Do Teams Keep Converting Points Into Hours Anyway?

Because hours feel more concrete, and most organizations report status in hours, dollars, and dates, not abstract point totals. A stakeholder asking "when will this ship" wants a date, and a PM under pressure will often do the conversion math privately: this sprint completed 32 points in 10 working days, so each point is roughly 0.3 days, so this 8-point story should take about 2.5 days. That conversion isn't wrong exactly, but it quietly reintroduces the exact bias story points were meant to route around: a promise about duration made before the work started, based on a rough average that assumes every story behaves like the average story.

The tell that a team has slipped back into hours-thinking is when someone starts debating whether a story is "really" a 5 or a 3 based on how many hours it feels like, instead of comparing it to the reference story. Once that debate starts, points have stopped functioning as a relative scale and have become a disguised hours estimate with extra steps. The fix isn't a rule against ever thinking about duration; it's noticing when the estimate conversation has quietly changed from "how does this compare to what we know" to "how many hours does this feel like," and pulling it back to the comparison.

What Velocity Is For, and What It Isn't

Velocity is for one thing: telling a specific team, using its own historical throughput, roughly how many points of already-estimated, already-ready work it can pull into the next sprint. That's it. It's an input to sprint planning, not an output that gets reported upward as a performance indicator. A team with a stable velocity of 30 points per sprint isn't a "better" team than one running at 18 points per sprint; they're two teams with different point calibrations, different team sizes, and different types of work, none of which velocity captures or was meant to capture.

What velocity isn't: a comparison tool between teams, a measure of individual output, a number that should trend permanently upward, or evidence of how hard anyone is working. A team's velocity legitimately drops when someone goes on leave, when a sprint absorbs more support tickets than usual, or when the team picks up unfamiliar, higher-risk work. None of those drops mean the team got worse at its job. They mean the team's actual available capacity changed, which is exactly the information velocity is supposed to surface so the next sprint gets planned against reality instead of last quarter's number.

Velocity Before and After Becoming a Target What happens when velocity becomes a target BEFORE: velocity tracked, not targeted S10 29 S11 31 S12 28 S13 30 Stable, roughly 28-31 pts/sprint AFTER: velocity reported as a KPI S14 33 S15 41 S16 48 S17 53 Points inflate; shipped features flat Same team, same underlying work rate, different label once someone starts watching the number

The diagram above shows the pattern that shows up almost every time an outside stakeholder starts watching a velocity chart. The left panel is a healthy team: stable, unremarkable, boring in the way a good metric should be boring. The right panel is the same team four sprints after someone above them started asking why velocity wasn't growing. The number climbs. What actually ships doesn't.

What Happens When Management Turns Velocity Into a Target?

Goodhart's Law describes it precisely: when a measure becomes a target, it stops being a good measure. Once velocity is a number that gets reported upward, reviewed in a leadership dashboard, or tied to a team's performance narrative, the team's incentive quietly shifts from "estimate accurately" to "produce a bigger number." Points inflate. A story that would have been called a 3 six months ago gets called a 5, not because anyone consciously decided to game the system, but because ambiguous estimation exercises drift toward whatever answer keeps the reviewing audience satisfied.

The damage compounds because the inflated number then gets used for exactly the planning purpose velocity was supposed to serve, and it fails at that job too. A team reporting 50 points per sprint that's actually delivering the same real output it delivered at 30 points will overcommit the next sprint, because the planning math now assumes a capacity that was never real. The team ends up worse off on both fronts: the metric is corrupted, and the thing the metric was supposed to help with, honest capacity planning, gets worse instead of better.

The fix isn't a policy against tracking velocity. It's keeping the number inside the team, where it's a planning tool, and refusing to let it become an externally reported performance indicator. A Scrum Master or team lead who protects the team from velocity becoming a KPI is doing exactly the job that role exists for: shielding the team's internal planning signals from becoming someone else's management theater.

How to Estimate in Points Without Anchoring on Time

  1. Pick a small, well-understood reference story first. Before estimating anything else, the team agrees on one story everyone has a shared, concrete understanding of, and assigns it a small point value, typically 1, 2, or 3.
  2. Compare every new story to the reference, not to a clock. The question is never "how many hours," it's "is this bigger, smaller, or about the same as our reference story, and by roughly how much."
  3. Use planning poker so estimates surface before anchoring happens. Every team member picks a number privately and reveals simultaneously. When two people land far apart, that gap is the valuable part of the exercise: it surfaces an assumption one person has that the other doesn't.
  4. Split anything the team can't confidently size. A story nobody can place relative to the reference usually isn't too big; it's too vague. Breaking it down until each piece is comparable to something already sized fixes the estimate, not a bigger number.
  5. Re-estimate at the reference level, not the individual-story level, when the team's makeup changes. If a new person joins or the team's skill mix shifts significantly, recalibrate the reference story's meaning as a team exercise rather than quietly letting new members import their old team's point scale.
  6. Let velocity settle over three to five sprints before trusting it. A single sprint's point total is noisy. The rolling average is what actually reflects the team's real, current capacity.

Why Comparing Velocity Across Teams Is Meaningless

Every team calibrates its point scale independently, starting from its own reference story. Team A might call a specific piece of backend refactoring work a 5. Team B, looking at the identical scope of work, might call it a 2, simply because Team B's reference story for "1 point" happened to be a bigger chunk of work than Team A's. There is no shared unit of measurement across teams the way there is with hours or dollars. A story point is meaningful only inside the team that generated it.

This is why leaderboards comparing team velocities are worse than useless: they don't just fail to measure anything real, they actively create the exact incentive problem described above, pushing every team toward point inflation to avoid looking slow next to a team whose points simply mean something different. If leadership needs a cross-team comparison, the honest options are comparing cycle time (a real, common unit of measurement, in days), throughput of completed items regardless of size, or outcome-level metrics like features shipped or customer-facing impact, none of which require pretending story points are portable across teams.

Reading Your Own Team's Velocity Trend Correctly

The useful signal in velocity isn't the absolute number, it's the trend and the variance. A team whose velocity has been stable within a narrow band for six sprints has a reliable planning number, whatever that number happens to be. A team whose velocity is swinging wildly, 20 points one sprint and 45 the next, has an estimation consistency problem worth investigating before the next planning cycle, regardless of what the average happens to land on.

Watch for a sustained downward trend that isn't explained by an obvious cause like reduced headcount or planned time off. That's the pattern worth a conversation, not because velocity dropping is inherently bad, but because it usually means something changed in the work itself, rising technical debt, more unplanned interruptions, growing scope ambiguity, that deserves attention on its own terms rather than being papered over by an artificially inflated point count. Onplana's sprint board tracks velocity per sprint automatically, computed from actual completed story points against elapsed sprint days, specifically so the trend is visible without anyone needing to reconstruct it from spreadsheets after the fact.

The Only Question Worth Asking About These Numbers

Every conversation about story points and velocity should come back to one test: is this number being used to help the team plan its own work, or is it being used to make the team look good or bad to someone outside it. The first use is exactly what the metric was built for. The second corrupts it within a couple of sprints, every time, because any number that becomes a target gets gamed, consciously or not.

Protect the boundary. Estimate in points because relative comparison is a genuinely better prediction tool than guessing hours. Track velocity because a team's own historical throughput is a genuinely useful planning input. Keep both inside the team's own planning process, and resist every request to turn either one into a scoreboard for people who were never going to use the number the way it was designed to be used.

Story PointsVelocityStory Points and VelocityAgile EstimationScrum MetricsSprint PlanningProject Management

Ready to make the switch?

Start your free Onplana account and import your existing projects in minutes.