
Development
Copilot on a Credit Budget
Table of Contents
A Quick Note on Timing
Everything here is accurate as of October 2026. The Copilot model list changes almost weekly and the pricing tables will drift. Check the current tables before you nerd rage at me.
I pulled the numbers from GitHub’s pricing pages and from a Microsoft Research paper, and I checked the math more than once, which still does not make it gospel. I am combining two sources that were never meant to meet, and rounding hard. If something looks wrong, tell me and I will fix it.
The Meter Moved
I have given up trying to keep pace with the Copilot model picker. Kimi K3 showed up in August, GPT-6 Astra in September, and as I write this GPT-6.1 Sol and Opus 5.5 are the newest things on the list, which will not be true by the time you read it. If your strategy for picking a model is “whatever is newest at the top of the dropdown” you do not have a strategy, you have a reflex.
What actually changed, and I think it matters more than any single model, is how GitHub bills for Copilot. For a long time you paid per prompt, and everything the agent did on its own after that was free. It was fucking awesome for consumers, but I can’t imagine that the economics made a lot of sense. The agent could grind on your task for an hour, make 40 tool calls, retry a broken build 6 times, and it cost you 1 request. That model is now legacy and the gravy train has left the station.
Now it is AI credits, across Chat, the CLI, the coding agent, Spark, all of it. Every token is on the meter, priced per model and converted at one credit per cent. Paddlin’ the school canoe? That is a credit. Pro includes 1,500 credits a month, Pro+ 7,000, Max 20,000, and past that you buy more at a cent each. In August I wrote that a few hundred dollars a month keeps 4 or 5 agents working for me at once, and I could not have told you where it went. Now I mostly can.
Which means the question is no longer “Which model is smartest”, but is “What does an agent session actually burn, and on which model is that burn worth it”. Until recently you could only guess at that, and now you do not have to.
What a Session Looks Like
In August a team at Microsoft Research published production traces from the Copilot coding agent: 13 million sessions, 3.2 million users, 761 million model calls, 95 trillion tokens, all from a single week in June. It is a systems paper about KV caches and GPU scheduling, but read as a customer it is an itemized bill.
Two words first. A turn is one prompt from you plus everything the agent does on its own before it comes back, and a session is the whole conversation.
The median session is 3 turns, 15 model calls, 13 tool calls, and 4 minutes, and the median turn sends about 228K tokens to the model, gets about 1.9K back, and roughly 217K of what it sent was a cache hit. Hold onto that last number, it is the whole ballgame.
The tail is brutal. By the 90th percentile a session is 15 turns and over 100 model calls, and the average session runs 15 times longer than the median one. Two more findings matter here. Tool failures hit 9% of turns and set off retry loops that multiply compute by up to 4, and switching models in the middle of a session drops the cache hit rate from about 90% to 8%.
The One Line That Outlives the Price List
Every model is priced on three numbers per million tokens, input, cached input, and output. Run the median turn through them and it collapses to this:
Credits per median turn ≈ input + (22 × cached input) + (0.2 × output)
Each term is that model’s dollar rate off GitHub’s pricing page, and you multiply by three for a median session. I checked it against every model in the picker and it lands within 1.5% of exact every time.
Take Sonnet 4.6, which bills $3.00 input, $0.30 cached, and $15.00 output. That is 3 + 6.6 + 3, about 12.6 credits a turn, roughly 37 or 38 a session. GPT-6 Luna bills $0.10, $0.01, and $0.50, and comes out to about 0.4 credits.
The part I keep coming back to is the middle coefficient. It weights the cached rate about 20 times heavier than the input rate, and here is why. On every call the agent re-sends the entire conversation, and almost all of it is unchanged from the call before. Both Anthropic and OpenAI serve that unchanged prefix from a cache and bill it at roughly 10% of the input rate, so on a workload that is 95% cache hits, the cached rate is usually the biggest single line in your bill and always the one with the most leverage. The headline input price that every model announcement leads with carries about a twentieth of the leverage, and I suspect that is not an accident.
So when the price list changes, pull the three numbers, run the line, and divide into whatever your plan gives you.
What It Cost on 6 October 2026
Here is that formula run across the picker as it stood on 6 October. I left GitHub’s own tier label in the second column on purpose, so you can watch it come apart. The model names are perishable, and the shape is the point. Cache-write charges are left out, which adds roughly 6% to 10% on Anthropic and the newer OpenAI models.
| Model | GitHub’s label | Per turn | Per session | Sessions a month on Pro |
|---|---|---|---|---|
| GPT-6 Luna | Lightweight | under 1 | ~1 | ~1,200 |
| GPT-5 mini | Lightweight | ~1 | ~4 | ~420 |
| Gemini 3.8 Flash (promo through Dec 31) | Versatile | ~3 | ~9 | ~160 |
| GPT-6.1 Sol | Powerful | ~6 | ~19 | ~80 |
| Claude Sonnet 5.5, GPT-6 Sol | Versatile / Powerful | ~8 | ~25 | ~60 |
| Claude Opus 5.5 | Powerful | ~12 | ~37 | ~41 |
| Claude Sonnet 4.6, Kimi K3 | Versatile / Powerful | ~12 | ~37 | ~40 |
| Claude Opus 5 | Powerful | ~21 | ~62 | ~24 |
| Claude Fable 5.1 | Powerful | ~25 | ~76 | ~20 |
| Claude Fable 5, GPT-6 Astra | Powerful | ~42 | ~125 | ~12 |
That is a 100x spread, on a workload where the cheap model often produces the identical diff.
Now look at the labels. GitHub files GPT-6.1 Sol, GPT-6 Sol, and Opus 5.5 as Powerful and Sonnet 4.6 as Versatile, and all three of the Powerful ones cost less per session than the Versatile one. Opus 5.5 undercuts Opus 5, listed right above it, by 40%, and Fable 5.1 beats Fable 5 by 40% on identical headline rates because only the cached rate changed. I do not think GitHub is being sneaky here, the labels describe capability and not cost, but the effect is the same and I would go by the pricing table instead.
One honest caveat. Those numbers multiply the median turn by the median session, and nobody’s real usage is median-shaped. Run the same math on the mean and a 37-credit session becomes about 205, which is 7 sessions a month on Pro instead of 40. The median is the task you finish in one clean pass, and the mean includes the afternoon spent arguing with a flaky integration test. Which one is you? Probably both, depending on the week.
How I Route Work Now
This is what I actually do, and it has changed since I wrote about Copilot in March. Back then Sonnet was my daily driver in VS Code and I reached for Opus to plan, and when the numbers moved, the habit moved with them.
- Inline completions are free. Completions and next edit suggestions are not billed in credits on any paid plan, so the cheapest agent session is the one I never start.
- The cheap end for the boring stuff. Unit tests for a class I just wrote, XML docs, converting a method to async. If the diff is easy to review, the cheap model is the right call.
- The 20 to 40 credit band for real work. The KumikoUI work I did this year, SkiaSharp cell renderers, edit triggers, the font registrar, is the kind of thing that lives here. When I ran the numbers that band held Opus 5.5 and GPT-6.1 Sol, so my default driver is a frontier model now.
- The expensive end for sessions I have already scoped. Cross-cutting refactors, unfamiliar code, bugs where the cause is not obvious. It earns its premium when the alternative is 3 failed sessions on a cheaper model, not on something I could describe in 2 sentences.
What the Traces Say to Stop Doing
Pick the model before the session, not during it. A mid-session switch drops the cache hit rate to 8%, so the next turn re-bills roughly 200K tokens at the full rate, and switching to a cheaper model halfway through can cost more than finishing on the one you started with. I imagine plenty of people are doing exactly that to save money.
Flaky tooling is now a line item, because failures hit 9% of turns and multiply compute by up to 4 through retries. While I was writing this a Hugo upgrade broke this blog’s Tailwind build for a few hours, and I kept thinking about what that would have cost if an agent had been the one retrying it.
And do not let sessions sprawl. Context compaction on long sessions throws away most of your prompt and resets the cache, so when a conversation starts to drift, a fresh scoped one is now the cheaper option and not just the tidier one.
If You Are Buying Seats
Team plans work differently, and it is mostly good news apart from a change on September 1 that quietly cost you 37% to 44% of your allowance.
Copilot Business includes 1,900 credits per user per month and Enterprise includes 3,900, pooled across the whole billing entity. 100 Business seats is a single pool of 190,000 credits, so your 2 developers who live in agent mode draw from the same pool as the 6 who mostly use inline completions, and the number to watch is pool burn rate rather than per-seat limits.
Here is the part that already bit people. From June 1 to September 1, existing Business and Enterprise customers had much higher included amounts, and that window has closed with nothing else changing. In fairness, GitHub documented the dates. I am just not sure anyone put them in a calendar, so here it is at Sonnet 4.6 rates:
| Plan | Credits per user | Median sessions | Mean sessions |
|---|---|---|---|
| Business, before Sept 1 | 3,000 | ~80 | ~15 |
| Business, today | 1,900 | ~51 | ~9 |
| Enterprise, before Sept 1 | 7,000 | ~187 | ~34 |
| Enterprise, today | 3,900 | ~104 | ~19 |
If you ran a Copilot pilot before September 1, that number is now wrong in the expensive direction. Rerun it, or divide the old figure by 1.6 for Business or 1.8 for Enterprise before anyone builds a budget on it, because the seats cost what they always cost and you get less for them now.
Measure Your Own Damn Team
The formula is built on somebody else’s median, and your team is not the median. Billing has an AI usage report with a per-user breakdown of credits consumed, capped at 31 days at a time, which is about a sprint.
So pick a sprint and have everyone work normally. At the end, pull the report, bucket consumption against what shipped, and put a number on your recurring task shapes. What does a service-layer refactor cost? A new endpoint with tests? A bug hunt? Write down pool burn as a percentage of the month, because that number survives a plan change.
Once you know a refactor runs, say, 150 credits on your default model, the plan conversation gets short. You know how many fit in a seat, who genuinely needs a bigger plan, and whether the expensive model earns 2x or 3x the credits on your codebase.
I do not know what the picker looks like in a month. I am fairly sure the three numbers will still be there.
Resources
Agentic Coding in the Wild: Characterizing GitHub Copilot at Production Scale (Microsoft Research, PDF) Models and pricing for GitHub Copilot Supported AI models in GitHub Copilot Billing reports reference Kimi K3 is now available in GitHub Copilot
Further Reading //


