The framework's sizing scale (XS through XL, in USD ranges) is a starting anchor based on observed behavior across a few teams I've worked with — not a benchmark. Per the calibration playbook, every team is expected to recalibrate to their own context.
For the framework to evolve from "one person's idea" to "useful empirical reference," we need real data from real teams.
What I'm looking for
Teams that have:
- Been using LLM-augmented development (Claude Code, Cursor, Aider, Copilot Workspace, etc.) for at least 2-3 months
- Tracked at least 20 tasks with both estimates and real costs
- Are willing to share anonymized aggregate numbers (no individual tasks, no per-developer data, no proprietary project info)
You don't have to use TokenPoints specifically — if you've been tracking inference cost per task in any form, your data is valuable.
What to share
The Share Calibration issue template covers it, but the minimum:
- Team size, language/stack, codebase size
- Model mix (% Sonnet / Opus / Haiku / other)
- Median real cost per size bucket
- Number of tasks tracked
- Any context that helps interpret the numbers
What you get back
- Your data (anonymized) goes into a public reference that helps every team that adopts this
- A credit in the README under "Contributors" (or anonymous if you prefer)
- Early access to v0.2 of the framework before it ships
What I'm explicitly NOT asking for
- Per-developer numbers
- Specific task descriptions
- Anything proprietary about your project, codebase, or business
If you're interested but unsure whether your data fits, comment here or DM me — happy to figure it out together.
The framework's sizing scale (XS through XL, in USD ranges) is a starting anchor based on observed behavior across a few teams I've worked with — not a benchmark. Per the calibration playbook, every team is expected to recalibrate to their own context.
For the framework to evolve from "one person's idea" to "useful empirical reference," we need real data from real teams.
What I'm looking for
Teams that have:
You don't have to use TokenPoints specifically — if you've been tracking inference cost per task in any form, your data is valuable.
What to share
The Share Calibration issue template covers it, but the minimum:
What you get back
What I'm explicitly NOT asking for
If you're interested but unsure whether your data fits, comment here or DM me — happy to figure it out together.