The desk is open
Same models.
Half the price.
By your deadline.
Offpeak — deadline-priced inference
30–60% of your AI spend has no human waiting — evals, embeddings, backfills, unattended agents. Give that work a deadline and it fills in batch tiers and off-peak windows at up to half price — SLA-guarded, on your own keys.
six venues, published
same model, computed live
that kept its batch
APACHE-2.0 · OPEN SOURCE · YOUR KEYS, PROVIDER-DIRECT · THE QUOTE IS FREE
01 // Quickstart
Half price in three steps
Install
Zero-dependency core; provider SDKs load only via extras. Python 3.10+, Apache-2.0, built in the open.
$ pip install "offpeak[all]" # or [openai] / [anthropic]
Give work a deadline
One argument on work you already run — batching, submission, polling, and result matching are handled. Your keys via the standard env vars; payloads go provider-direct and never transit Offpeak. If a batch runs long, the stragglers re-run synchronously at list: you state the deadline; it gets met.
import offpeak jobs = [offpeak.job("claude-haiku-4-5", f"Summarize:\n\n{doc}") for doc in docs] results = offpeak.run(jobs, deadline="06:00") # done by 6am, batch prices print(offpeak.receipt(results))
Settle the receipt
Every run prints one — list, paid, captured, arithmetic against the published sheet. This one is from a real settled run on Google's batch tier: 50.0% captured, 5/5 SLAs, zero fallbacks. The ledger holds every receipt.
OFFPEAK SETTLEMENT ──────────────────────────── jobs 5 (5 ok, 0 sync fallback) sla 5/5 met venues gemini:batch 5 tokens 124 in · 2,427 out list $0.00919 paid $0.00460 captured $0.00460 (50.0%) prices snapshot 2026-08-30 — offpeak.prices ───────────────────────────────────────────────
No key yet? The quote is free — python -m offpeak quote --model gpt-5.6-luna --jobs 5000 … prices a run against the published sheets with no API calls and no key at all.
02 // How it works
One argument
Say it can wait
Work you already run gains one argument. Nothing else changes. Urgent traffic is never touched.
offpeak.run(jobs, deadline="06:00")The desk places it
Portfolio-scheduled across batch tiers, off-peak windows and clock-priced lanes — a fallback venue guards every deadline. Your runner, your keys: payloads never transit the desk.
Settle the receipt
At the deadline: spread captured, SLA met, energy logged. Every dollar checkable against a public price sheet.
captured $10.56 · 50.0% · SLA 6,000/6,00003 // The spread
The price of urgency
Exhibit 01 // published sheets · 2026-08-30
The spread is old. Batch tiers at half of list have sat on published sheets for years, and day-ahead power markets have priced the same valley for decades — same model, same tokens, the only difference is when it lands. Venues now price urgency in explicit tiers, fastest to batch 4× apart. The premium is a published number, not our estimate.
- 4× fastest ↔ batch tier
- 50% off batch — six venues
- 80% off clock-promo lane
- $0 modeled — all printed
Full brief
04 // The book
The open book
Exhibit 02 // deadline contracts · 6,000 rows · showing 5
- 6,000 contracts · 12.68M tokens
- $21.12 list — the run-it-now price
- $10.56 projected paid · 50.0% captured
- $23.24 hard ceiling, all-sync worst case
| Contract | Model | Venue | Deadline | Status |
|---|---|---|---|---|
| EMB-BACKFILL ×1,000EMBEDDINGS | sonnet-5 | anthropic batch | 06:00 | FILLED |
| EVAL-SUITE ×800REGRESSION | haiku-4-5 | anthropic batch | 05:00 | FILLED |
| EVAL-LONGTAIL ×120SLOW EVALS | terra | openai → own-GPU 02:30 | 04:30 | RE-ROUTED |
| CLASSIFY ×2,000TICKET TRIAGE | luna | openai batch | 06:00 | WORKING |
| URGENT CONTROL ×80INTERACTIVE — NEVER TOUCHED | sonnet-5 | spot · standard | — | SPOT |
Full brief
05 // The valley
Into the valley
Exhibit 03 // known, not forecast
The same kilowatt-hour, 1.74× apart inside 24 hours. Day-ahead clears before delivery — the next period's prices are knowledge, not forecast. The desk schedules into the valley and holds a fallback venue for every deadline.
- 1.43× GB · 1.27× CAISO
- 3.94× ERCOT, same 24 hours
Full brief
06 // The receipt
Settled, on time
DEMO BOOK · REAL CURVES. The dollar figures above are a simulated 6,000-job book.
REAL SETTLEMENTS EXIST. Every run that reached a batch tier and kept it settled at exactly 50.0% of list — real bills, four capturing venues, every receipt public and checkable against the published sheet. The runs that captured nothing — a gated venue, and one night offpeak's own driver discarded a batch that had already succeeded — are on the same ledger. The ledger.
Every dollar reproducible from a public price sheet. Quotes ≠ settlements — settlements only ever come from real runs.
What waiting is worth
Savings calculator
Estimate your monthly spread
Two ways in. Start from a monthly bill and a guess at how much of it has no human waiting — or build the book workload by workload against the real published sheet. Both price the same thing: the urgency premium you are paying today.
01 // Estimate
What a deadline is worth
You would stop paying
$6,800
per month · 17.0% off the whole bill
Every number here is arithmetic on a published price sheet — no modeled inputs, no assumed energy, nothing that needs a footnote. Actual capture depends on how much of your work really has slack.
| Workload | Model | Requests / mo | Tokens in | Tokens out | Paying now | Deadline |
|---|
You would stop paying
$0
per month · 0% off this book
Prices are the sheet the SDK ships — offpeak.prices, snapshot 2026-08-30, re-verified against the rendered provider pages 2026-08-30, reproduced in 02 // The sheet below. Batch is exactly half of standard on every venue — on DeepSeek the same half is the off-peak clock rate, not a batch queue; fast tiers are published on the gpt-5.6 family and, as a research preview, on Claude Opus 5 and Opus 4.8. Loaded by default: the real parked book — 6,000 jobs, 12.68M tokens, $28.05 list.
02 // The sheet
What the prices are
offpeak.prices // snapshot 2026-08-30 · USD per 1M tokens
| Model | Standard in | Standard out | Batch / off-peak in | Batch / off-peak out | Fast in | Fast out |
|---|
gpt-5.6-sol's standard rate is promotional at least through 2026-11-21; post-promo list is $5 / $30 and both tiers move with it, so the ratio is the durable figure, not the dollars. Batch is exactly 50% of standard. The openai/gpt-oss rows are Groq's, the mistral rows are from mistral.ai/pricing/api, and the gemini rows from ai.google.dev/pricing — all three at the same 50% batch rule. Groq and Mistral publish no per-model fast rate, so this sheet carries none; Google prices a priority tier that this sheet does not carry yet, and tools/sheet_reconcile.py reports the gap rather than the sheet implying a spread it has not verified. Gemini 3.7 and 3.6 Flash are introductory through 2026-12-31 and then double to $1.50 / $7.50 on both legs, which moves the dollars and leaves the batch ratio alone. The deepseek rows (api-docs.deepseek.com) carry the peak rate as standard; DeepSeek has no batch API, and its half price is the off-peak clock — peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays, everything else and all weekend is half — so the discount is decided per request by the wall clock, not by a queue. The qwen rows are Alibaba Model Studio's international region, batch at 50%; qwen3.7-max is "limited-time 50% off" with no published end date, so it carries no expiry and may step up unannounced. Fast tiers are published on the gpt-5.6 family and on Claude Opus 5 / Opus 4.8 ($10 / $50, research preview, first-party API only, not combinable with batch) — fastest to batch is 4× at both venues, computed live by urgency_spread(), never cached. Sonnet 5's $2 / $10 is now Anthropic's standard price; the scheduled 09-01 rise to $3 / $15 will not occur.
03 // Honest caveats
What this does not prove
Not a quote
Actual capture depends on what share of your work really has slack, and on rates holding. A quote comes from your own metadata, free, in the flexibility report.
Insurance costs
A job rescued at its deadline pays list, not batch. Smoke runs to date: zero fallbacks, 6/6 SLAs — but the model above assumes none, and real books have some.
Quoting daily
Spread Board
Offpeak's public track record
What waiting is worth, in public: the three spreads a deadline can capture — tokens, power, carbon — quoted from open data at 19:00Z, marked against actuals at 06:30Z, daily, against the deadline and not the clock.
batch at half of list
day-ahead, this session
chosen window
QUOTES COST NOTHING · SETTLEMENTS ONLY EVER COME FROM REAL RUNS
01 // The live quote
Three legs, latest session
Tokens // both venues
Batch at 50% of list on both published sheets. Fast → batch, at both venues that sell one: 4×.
Power // GB · Agile day-ahead
34.85p across the 17–21 peak → 25.23p in the 00–05 window. Extremes 40.99p / 23.61p.
Carbon // GB · NESO forecast
Cleanest 5h 110.9g · dirtiest 159.5g. The naive 00–05 shift is 0.74× — worse than peak.
Full brief
02 // Marked
Marked against actuals
Marked sessions // actuals
| Leg · zone | High · latest | Low · latest | Prior session | Latest marked | Note |
|---|---|---|---|---|---|
| POWER · GBAGILE, DAY-AHEAD | 37.17p | 24.82p | 1.43× | 1.50× | extremes 1.74× and 2.28× |
| POWER · CAISO SP15DA LMP $/MWh | $82.52 | $50.06 | 1.27× | 1.65× | evening peak, solar rolled off |
| POWER · ERCOT HOUDA LMP $/MWh | $93.83 | $23.70 | 3.94× | 3.96× | widest zone on the board, twice |
| CARBON · GBNESO · CHOSEN WINDOW | 146.1 g | 65.0 g | 1.13× | 2.25× | naive 00–05 reads 1.03× and 2.10× |
| CARBON · CISO + ERCOEIA-930 · DERIVED · BACKFILLED | — | — | 0.80× / 0.80× | — | both sub-parity; the prior session awaits its following-morning trough |
| TOKENS · OAI + ANTPUBLISHED SHEETS | list | 50% list | 2.0× | 2.0× | constant until a vendor moves a sheet — which this board catches |
Methodology & sources
03 // The carbon leg
Cheap hours are not clean hours
Exhibit 03 // GB carbon · every session marked
The board's third leg is the one that misbehaves. Run the naive version — "shift it to 00:00–05:00" — and the carbon result is a coin flip: 1.03×, 2.10×, and 0.74×. That last one means the window was dirtier than the evening peak. Evening solar is gone by midnight and what backfills it is not clean. CAISO has done the same, at 0.76×.
Pick the window from the curve instead of assuming it and the same sessions read 1.13×, 2.25× and 1.44× — never below parity. Price and carbon are two different curves and get optimized separately.
Definitions & sources
Anything left of parity is a session where shifting the work would have added carbon. The chosen window has never landed there.
NESO · keyless · marked 06:30Z04 // Settlements
Real runs only
The settled record // real money, real venues · SETTLED.md
| Run | Venues | Jobs | List | Paid | Captured | SLA |
|---|---|---|---|---|---|---|
| mechanics-1TWO VENUES · OPENAI LEG REJECTED | anthropic + openai batch | 48 | $0.00156 | $0.000781 | 50.0% | 24/48 |
| mechanics-2CEILING TOO LOW · EMPTY OUTPUTS | openai batch | 24 | $0.000594 | $0.000297 | 50.0% | 24/24 |
| mechanics-3CEILING SIZED TO THE MODEL | openai batch | 24 | $0.00131 | $0.000655 | 50.0% | 24/24 |
| groq-1BATCH TIER GATED · SYNC FALLBACK AT LIST | groq · sync fallback | 24 | $0.00236 | $0.00236 | 0.0% | 24/24 |
| mistral-1BATCH TIER GATED · SYNC FALLBACK AT LIST | mistral · sync fallback | 24 | $0.000199 | $0.000199 | 0.0% | 24/24 |
| gemini-1BATCH TIER · FIRST THIRD VENUE TO CAPTURE | gemini batch | 5 | $0.00919 | $0.00460 | 50.0% | 5/5 |
| 08-26 openai-1BATCH TIER | openai batch | 24 | $0.00127 | $0.000636 | 50.0% | 24/24 |
| 08-26 anthropic-1BATCH TIER | anthropic batch | 6 | $0.000378 | $0.000189 | 50.0% | 6/6 |
| 08-26 gemini-1BATCH TIER · REPEATED TWO DAYS ON | gemini batch | 5 | $0.00918 | $0.00459 | 50.0% | 5/5 |
| 08-26 mistral-2BATCH COMPLETED, THEN DISCARDED · CLIENT BUG | mistral · sync fallback | 24 | $0.000198 | $0.000198 | 0.0% | 24/24 |
| 08-26 mistral-3BATCH TIER · FOURTH VENUE TO CAPTURE | mistral batch | 24 | $0.000200 | $0.0000999 | 50.0% | 24/24 |
Scale is on every row on purpose: these are mechanics proofs, not production volume — the evidence they carry is that the submit → poll → settle path works and prices out at exactly half of list on a real bill. Three rows captured nothing, and they stay. Groq (403 not_available_for_plan) still gates the batch API behind a plan, so those jobs took the sync fallback and paid list. Mistral was gated the same way (402) until billing was switched on — and then failed a second time for a different reason: on 08-26 the batch completed and was thrown away by a bug in offpeak's own driver, which reached for a streaming response's text before reading it. That run paid twice for the same answers and captured nothing. It is on the board next to the fixed re-run, because a client that loses a batch costs exactly what a venue that refuses one costs, and is harder to notice. The full record.
Real money, published whole
Settled runs
The settled record
Real runs on real bills, across five venues — four of them now capturing: OpenAI, Anthropic, Google Gemini and Mistral. Every run that reached a batch tier and kept it captured exactly 50.0% against the published sheet. Runs that captured nothing are published anyway — a gated venue, and one night where the batch completed and offpeak's own driver threw it away. The failures are the reason to believe the rest. Scale is printed on every row on purpose: this ledger proves the path, not the volume.
that kept its batch
at the batch tier
rather than dropped
01 // The record
The record, whole
Receipts // generated by tools/mechanics_run.py + settle_report.py · SETTLED.md
| Run | Venues | Jobs | Tokens | List | Paid | Captured | SLA |
|---|---|---|---|---|---|---|---|
| mechanics-148 JOBS, TWO VENUES | anthropic 24 · openai 24 | 48 | 787 / 155 | $0.00156 | $0.000781 | 50.0% | 24/48 |
| mechanics-2CEILING TOO LOW | openai batch | 24 | 728 / 374 | $0.000594 | $0.000297 | 50.0% | 24/24 |
| mechanics-3CEILING SIZED TO THE MODEL | openai batch | 24 | 728 / 971 | $0.00131 | $0.000655 | 50.0% | 24/24 |
| groq-1BATCH TIER GATED · SYNC FALLBACK AT LIST | groq · sync fallback | 24 | 2,288 / 7,307 | $0.00236 | $0.00236 | 0.0% | 24/24 |
| mistral-1BATCH TIER GATED · SYNC FALLBACK AT LIST | mistral · sync fallback | 24 | 948 / 94 | $0.000199 | $0.000199 | 0.0% | 24/24 |
| gemini-1BATCH TIER · FIRST THIRD VENUE TO CAPTURE | gemini batch | 5 | 124 / 2,427 | $0.00919 | $0.00460 | 50.0% | 5/5 |
| 08-26 openai-1BATCH TIER · 8m07s | openai batch | 24 | 728 / 938 | $0.00127 | $0.000636 | 50.0% | 24/24 |
| 08-26 anthropic-1BATCH TIER · 4m05s | anthropic batch | 6 | 188 / 38 | $0.000378 | $0.000189 | 50.0% | 6/6 |
| 08-26 gemini-1BATCH TIER · 2m04s | gemini batch | 5 | 124 / 2,422 | $0.00918 | $0.00459 | 50.0% | 5/5 |
| 08-26 mistral-2BATCH COMPLETED, THEN DISCARDED · CLIENT BUG | mistral · sync fallback | 24 | 948 / 93 | $0.000198 | $0.000198 | 0.0% | 24/24 |
| 08-26 mistral-3BATCH TIER · FOURTH VENUE TO CAPTURE · 33s | mistral batch | 24 | 948 / 96 | $0.000200 | $0.0000999 | 50.0% | 24/24 |
Tokens shown as in / out. Every figure is emitted by the settlement code from counted tokens against the published sheet — none of it is typed in by hand, and the arithmetic reproduces in the calculator.
02 // What went wrong
Five failures, published
The venue said no
Anthropic settled 24/24. OpenAI rejected all 24 with an HTTP 400 — the driver sent a token-ceiling parameter the gpt-5.6 family no longer accepts. Nothing was billed on the failed leg.
SLA 24/48 · $0 on the rejected venueBilled for nothing useful
The OpenAI leg re-ran and settled 24/24 — and returned 24 empty strings, because a 16-token ceiling was spent entirely on reasoning. Billed, SLA met, output worthless. Published rather than dropped.
SLA 24/24 · 374 output tokens · 0 answersSame lines, right ceiling
The same 24 lines at a 256-token ceiling: 24/24 real answers, 971 output tokens, zero fallbacks, every deadline met, half of list paid.
SLA 24/24 · captured $0.000655 · 50.0%The venue was gated
Groq's entire batch surface answered 403 not_available_for_plan — file upload and batch listing alike — so submission never completed and all 24 jobs took the sync fallback at list price. Billed, every deadline met, nothing captured. The blocker is plan entitlement, not the driver.
SLA 24/24 · $0.00236 paid · 0.0% capturedAnd the probe missed it
Mistral answered 402 · enable billing via the console on batch create. Worse, an entitlement probe had cleared it a day earlier: a malformed request came back with a field-level validation error, so it looked reachable. Validation runs before the billing check — a request that cannot succeed can never reach the paywall.
SLA 24/24 · $0.000199 paid · 0.0% capturedGemini captures
Google Gemini batch, once billing was enabled: 5/5, zero sync fallbacks, 50.0% captured, 2m36s end to end. The first venue here beyond OpenAI and Anthropic to collect the spread rather than record why it could not.
SLA 5/5 · captured $0.00460 · 50.0%The batch worked. We didn't.
Mistral's paywall was gone and the batch succeeded — 24/24 in 6m16s. offpeak threw it away: _download reached for a streaming response's .text before reading it, which raises ResponseNotRead — a RuntimeError, so the getattr(…, None) guard never applied. run() booked a finished batch as a polling failure and paid list for all 24. It had completed 32 seconds earlier.
Mistral captures
Same twenty-four lines, ninety minutes later, on the fixed driver: 24/24, zero fallbacks, 50.0% captured, 33 seconds end to end. Nothing about the venue changed — only whether offpeak could read the file it was already being handed.
SLA 24/24 · captured $0.0000999 · 50.0%03 // The parked book
The 6,000-job exhibit
Re-priced against current published rates, the full demonstration book came to $28.05 list against a $23.24 authorized ceiling — over, so it was never built and never submitted. It stays parked until a partner conversation needs the exhibit, or until a partner's own runner settles it on their keys, which costs Offpeak nothing.
| Job class | Model · venue | Jobs | Tokens / job | List $ | Batch $ |
|---|---|---|---|---|---|
| Summarize 5k-token doc chunksANTHROPIC BATCH | claude-sonnet-5 | 1,000 | 5,000 / 300 | 13.00 | 6.50 |
| Summarize 5k-token doc chunksOPENAI BATCH | gpt-5.6-terra | 1,000 | 5,000 / 300 | 13.60 | 6.80 |
| Classify short passagesANTHROPIC BATCH | claude-haiku-4-5 | 2,000 | 500 / 20 | 1.20 | 0.60 |
| Classify short passagesOPENAI BATCH | gpt-5.6-luna | 2,000 | 500 / 20 | 0.25 | 0.12 |
| Re-priced at current rates~12.7M TOKENS | — | 6,000 | 28.05 | 14.02 |
03 // Honest caveats
What this record does not prove
Sub-cent is sub-cent
The captured dollars are deliberately tiny. What is proved is the path and the arithmetic, on real bills. Production volume is a different claim and is not made here.
Offpeak's own work
These ran on Offpeak's keys against Offpeak's own book. No buyer receipt is on this ledger yet — that is the open half of the proof, and the point of the design-partner program.
Insurance costs
A job rescued at its deadline pays list, not batch. Zero fallbacks on the clean runs — but run 1 showed the fallback has a blind spot, and a real book will find more.
Sheets watched daily
Prices
Every number the desk quotes against, drawn
Three series, one page: the token sheet every venue publishes and the desk prices against, the power marks that show the same valley in three grids, and the batch turnaround our canaries measure — the number no sheet prints.
same model, same tokens
peak over off-peak, latest marked
median, completed canaries
Sheet snapshot · board-data as committed
01 // The sheet
Same model, three prices
Every row is one model at one venue. The dim end is the batch tier — or, on DeepSeek, the off-peak clock rate, half of peak by time of day rather than by queue — the bright end is standard, and where a venue sells a fast tier it sits further right. The gap is what waiting is worth — log scale, because the sheet spans three orders of magnitude and the ratio is the durable figure, not the dollars.
Batch → standard → fastoutput · $/M tokens
SOURCE: offpeak.prices, PROVIDER PUBLIC SHEETS, SNAPSHOT 2026-08-30, RE-VERIFIED FROM RENDERED PAGES. BATCH = EXACTLY HALF OF STANDARD ON EVERY ROW SHOWN. DEEPSEEK SELLS NO BATCH: ITS STANDARD ROW IS THE PEAK RATE (01:00–04:00 + 06:00–10:00 UTC, MON–FRI) AND THE DIM END IS THE OFF-PEAK RATE, HALF OF PEAK BY THE WALL CLOCK — SAME ARITHMETIC, DIFFERENT LANE. QWEN (ALIBABA MODEL STUDIO, INTERNATIONAL REGION) BATCHES AT 50%; QWEN3.7-MAX IS "LIMITED-TIME 50% OFF" WITH NO PUBLISHED END DATE. FAST = 2× STANDARD WHERE SOLD: THE GPT-5.6 FAMILY, AND CLAUDE OPUS 5 / OPUS 4.8 AS A RESEARCH PREVIEW (FIRST-PARTY API ONLY, NOT WITH BATCH) — 4× FAST → BATCH AT BOTH VENUES. SHEET RE-VERIFIED AGAINST THE RENDERED PROVIDER PAGES 2026-08-30. PROMOTIONAL RATES (SOL ≥ 2026-11-21 · GEMINI 3.6/3.7 FLASH ≤ 2026-12-31) CARRY THEIR EXPIRY AND RE-DERIVE, NEVER CACHED. CACHED-INPUT AND LONG-CONTEXT COLUMNS ARE NOT DRAWN.
02 // The marks
The same valley, three grids
Power is the leg with a public day-ahead price. Each session the board quotes the 17:00–21:00 local evening peak against the 00:00–05:00 trough and marks it the next morning. The ratio panel puts all three grids on one axis; the level panels show the dollars behind it, in each grid's own unit.
Peak over off-peakratio · marked at 06:30Z
GBp/kWh
CAISO$/MWh
ERCOT$/MWh
SOURCES: GB — OCTOPUS AGILE DAY-AHEAD (KEYLESS, REGION C), HALF-HOURLY. CAISO SP15 AND ERCOT HOUSTON — DAY-AHEAD LMP VIA gridstatus, HOURLY. A SESSION SPANS 16:00Z–07:00Z; THE OFF-PEAK WINDOW LANDS THE FOLLOWING LOCAL MORNING, SO THE US LEGS MARK ONE DAY BEHIND GB. GAPS ARE SESSIONS WHOSE TROUGH HOURS HAD NOT LANDED WHEN THE MARK WAS TAKEN — RECORDED AS UNAVAILABLE, NEVER FILLED IN. EACH POINT IS ONE nightly/<session>-mark.json ON BOARD-DATA; RELATIVE LABELS ON THE AXIS, THE SESSION DATE ON HOVER AND IN THE TABLE.
03 // The queue
How long the cheap lane takes
No venue publishes this. Every session the desk submits a two-job canary into each venue's batch tier with a 24-hour window and records how long it actually took. Most clear in minutes; one took eight and a half hours. A hollow marker at the top is a censored leg — still running when the session closed, so all we know is "longer than we watched."
Coverage // what is measured, and what is not · as of 2026-08-30
The chart below draws only venues with a measured queue — four today. DeepSeek's clock lane is measured differently: there is no batch to wait for, so the probe records which regime the clock is in and the synchronous latency of a request made in it. Hourly rows stay in the private probe repo; this page carries the daily series, per-venue percentiles and counts — never hour-of-day.
| Venue | Lane | Discount | Status | Since | Note |
|---|---|---|---|---|---|
| OpenAIBATCH API | batch | 50% off | MEASURED | daily 08-23 · hourly 08-30 | Capturing. Two-job canary, 24 h window; the daily series is public, the hourly rows are private. |
| AnthropicMESSAGE BATCHES | batch | 50% off | MEASURED | daily 08-23 · hourly 08-30 | Capturing. Same canary, same window. |
| Google GeminiBATCH MODE | batch | 50% off | MEASURED | daily 08-26 · hourly 08-30 | Capturing. Same canary, same window. |
| MistralBATCH API | batch | 50% off | MEASURED | daily 08-26 · hourly 08-30 | Capturing. Same canary; the one eight-and-a-half-hour turnaround on the chart is this venue's. |
| DeepSeekCLOCK LANE · NO BATCH API | clock | 50% off-peak | MEASURED | hourly 08-30 | Not a queue, so not on the chart. Observable is regime plus synchronous latency: first two observations 2.3 s and 2.6 s, off-peak, 2,000 / 300 tokens. |
| QwenALIBABA MODEL STUDIO · INTL | batch | 50% off | PROBATION | first live run 08-30 | Driver shipped in 0.2.7. Joins the chart once it has settled repeatedly. |
| GroqBATCH API | batch | 50% off | BLOCKED | — | Batch tier answers 403 — plan-gated, and plan sales are suspended. Sync fallback only. |
| xAIFLAGSHIP | — | −20% legacy only | NO LANE | — | No batch lane on the flagship; the discount is on legacy models only. Watched, not probed. |
| AWS BedrockBATCH INFERENCE | batch | 50% off | NOT BUILT | — | Published; driver on partner pull. |
| AzureBATCH | batch | — | NOT BUILT | — | On partner pull. |
| KimiMOONSHOT | batch | 40% off | NOT BUILT | — | Published at 40%; not built. |
Turnaround by venuelog scale · 24 h window
Inline series · the floor
Per-venue summarypercentiles · counts
SOURCE: THE DAILY SERIES ABOVE IS nightly/<session>-queue.json ON BOARD-DATA AS COMMITTED, WRITTEN BY tools/queue_probe.py; THE PROBE NOW RUNS HOURLY FROM A PRIVATE REPO AND PUBLISHES BACK nightly/QUEUE.md AND nightly/queue-summary.json — PER-VENUE PERCENTILES AND COUNTS, NEVER HOUR-OF-DAY, NEVER A PER-ROW HOURLY SERIES. TWO JOBS PER LEG, ~30 TOKENS IN, SUB-CENT SPEND CAPPED AT $0.01 PER SESSION. SKIPPED LEGS (NO KEY IN THE ENVIRONMENT) ARE NOT DRAWN AND STAY IN THE TABLE. THIS IS THE REALIZED HALF OF THE RECORD — THE SHEET SAYS "WITHIN 24 HOURS"; THIS SAYS WHAT THAT MEANT.