The desk is open

Same models.
Half the price.
By your deadline.

Offpeak — deadline-priced inference

30–60% of your AI spend has no human waiting — evals, embeddings, backfills, unattended agents. Give that work a deadline and it fills in batch tiers and off-peak windows at up to half price — SLA-guarded, on your own keys.

50% off
Batch tier
six venues, published
Fastest → batch
same model, computed live
50.0%
Captured on every run
that kept its batch
Quickstart Estimate savings

APACHE-2.0 · OPEN SOURCE · YOUR KEYS, PROVIDER-DIRECT · THE QUOTE IS FREE

01 // Quickstart

Half price in three steps

STEP 01

Install

Zero-dependency core; provider SDKs load only via extras. Python 3.10+, Apache-2.0, built in the open.

shell
$ pip install "offpeak[all]"   # or [openai] / [anthropic]
STEP 02

Give work a deadline

One argument on work you already run — batching, submission, polling, and result matching are handled. Your keys via the standard env vars; payloads go provider-direct and never transit Offpeak. If a batch runs long, the stragglers re-run synchronously at list: you state the deadline; it gets met.

batch_summarize.py
import offpeak

jobs = [offpeak.job("claude-haiku-4-5", f"Summarize:\n\n{doc}")
        for doc in docs]

results = offpeak.run(jobs, deadline="06:00")  # done by 6am, batch prices
print(offpeak.receipt(results))
STEP 03

Settle the receipt

Every run prints one — list, paid, captured, arithmetic against the published sheet. This one is from a real settled run on Google's batch tier: 50.0% captured, 5/5 SLAs, zero fallbacks. The ledger holds every receipt.

receipt · 2026-08-24-gemini-1 · real run
OFFPEAK SETTLEMENT ────────────────────────────
jobs      5 (5 ok, 0 sync fallback)
sla       5/5 met
venues    gemini:batch 5
tokens    124 in · 2,427 out
list      $0.00919
paid      $0.00460
captured  $0.00460 (50.0%)
prices    snapshot 2026-08-30 — offpeak.prices
───────────────────────────────────────────────

No key yet? The quote is freepython -m offpeak quote --model gpt-5.6-luna --jobs 5000 … prices a run against the published sheets with no API calls and no key at all.

02 // How it works

One argument

STEP 01

Say it can wait

Work you already run gains one argument. Nothing else changes. Urgent traffic is never touched.

offpeak.run(jobs, deadline="06:00")
STEP 02

The desk places it

Portfolio-scheduled across batch tiers, off-peak windows and clock-priced lanes — a fallback venue guards every deadline. Your runner, your keys: payloads never transit the desk.

STEP 03

Settle the receipt

At the deadline: spread captured, SLA met, energy logged. Every dollar checkable against a public price sheet.

captured $10.56 · 50.0% · SLA 6,000/6,000

03 // The spread

The price of urgency

Exhibit 01 // published sheets · 2026-08-30

The spread is old. Batch tiers at half of list have sat on published sheets for years, and day-ahead power markets have priced the same valley for decades — same model, same tokens, the only difference is when it lands. Venues now price urgency in explicit tiers, fastest to batch 4× apart. The premium is a published number, not our estimate.

  • fastest ↔ batch tier
  • 50% off batch — six venues
  • 80% off clock-promo lane
  • $0 modeled — all printed
Full brief
BATCH AT HALF OF LIST, PUBLISHED: OAI · ANTHROPIC · GOOGLE · MISTRAL · GROQ · BEDROCK — THE STRUCTURAL SPREAD, STABLE UNTIL A VENDOR MOVES A SHEET. THE 4× IS THE CURRENT FASTEST-TO-BATCH RANGE AT OPENAI AND, IN PREVIEW, AT ANTHROPIC, COMPUTED LIVE VIA urgency_spread() AGAINST WHATEVER THE SHEET SAYS TODAY — PART OF IT RIDES ON PROMOTIONAL PRICING AND RE-DERIVES WHEN THAT LAPSES. CLOCK-PRICED LANES (80% OFF OFF-PEAK; ANOTHER ≈0.5× PEAK) ARE REGIONAL AND TIME-BOUND. NO SPREAD IS QUOTED FROM A STALE SHEET — DECAYING RATES ARE RE-VERIFIED BEFORE ANY CUSTOMER-FACING MATH.
Evidence 01 · multiple of spot 2.0× 1.0× 0.5× 0.2× FAST 2.00 SPOT 1.00 BATCH 0.50 PROMO 0.20 NOW≤ 24 HCLOCK Uplink 0x551A · live sheets

04 // The book

The open book

Exhibit 02 // deadline contracts · 6,000 rows · showing 5

  • 6,000 contracts · 12.68M tokens
  • $21.12 list — the run-it-now price
  • $10.56 projected paid · 50.0% captured
  • $23.24 hard ceiling, all-sync worst case
ContractModelVenueDeadlineStatus
EMB-BACKFILL ×1,000EMBEDDINGSsonnet-5anthropic batch06:00FILLED
EVAL-SUITE ×800REGRESSIONhaiku-4-5anthropic batch05:00FILLED
EVAL-LONGTAIL ×120SLOW EVALSterraopenai → own-GPU 02:3004:30RE-ROUTED
CLASSIFY ×2,000TICKET TRIAGElunaopenai batch06:00WORKING
URGENT CONTROL ×80INTERACTIVE — NEVER TOUCHEDsonnet-5spot · standardSPOT
Full brief
FULL BOOK: SUMM-CORPUS ×700 (TERRA, OAI BATCH) · AGENT-MAINT ×900 + RERANK ×400 ON OWN-GPU vLLM 00–05 · TENORS T+6.5H → T+8H. CONTROL PLANE ONLY — PAYLOADS NEVER TRANSIT THE DESK; THE CUSTOMER'S RUNNER EXECUTES ON THEIR KEYS. URGENT TRAFFIC IS NEVER TOUCHED. QUOTES FROM OPEN DATA COST NOTHING; SETTLEMENTS ONLY EVER COME FROM REAL RUNS.

05 // The valley

Into the valley

Evidence 03 · GB power, day-ahead · p/kWh 41.6p 24.0p 16:0000–0506:30 Agile day-ahead · real marked session

Exhibit 03 // known, not forecast

The same kilowatt-hour, 1.74× apart inside 24 hours. Day-ahead clears before delivery — the next period's prices are knowledge, not forecast. The desk schedules into the valley and holds a fallback venue for every deadline.

  • 1.43× GB · 1.27× CAISO
  • 3.94× ERCOT, same 24 hours
01:38:52QUEUE P95 BREACHre-route 120 → own-GPU · SLA HELD
Full brief
MARKS, SESSION 08-20→21 (POWER, DAY-AHEAD): GB 1.43× · CAISO 1.27× · ERCOT 3.94×. CARBON GB 1.03× — NEAR-FLAT, REPORTED AS-IS. RISK: P95 QUEUE LATENCY VS BUFFER PER VENUE; A BREACH TRIGGERS RE-ROUTE TO THE OWN-GPU WINDOW (00–05, 24.7p/kWh, 8×4090). A RESCUED JOB PAYS LIST — IT NEVER BREACHES.

06 // The receipt

Settled, on time

Captured
$4.9104
Filled
2,786 / 6,000
SLA
100%
Energy
2.1 kWh

DEMO BOOK · REAL CURVES. The dollar figures above are a simulated 6,000-job book.
REAL SETTLEMENTS EXIST. Every run that reached a batch tier and kept it settled at exactly 50.0% of list — real bills, four capturing venues, every receipt public and checkable against the published sheet. The runs that captured nothing — a gated venue, and one night offpeak's own driver discarded a batch that had already succeeded — are on the same ledger. The ledger.

Every dollar reproducible from a public price sheet. Quotes ≠ settlements — settlements only ever come from real runs.

What waiting is worth

Savings calculator

Estimate your monthly spread

Two ways in. Start from a monthly bill and a guess at how much of it has no human waiting — or build the book workload by workload against the real published sheet. Both price the same thing: the urgency premium you are paying today.

01 // Estimate

What a deadline is worth

Your stack
$40,000
Provider API bills only — not seats, not infra
40%
Evals, embeddings, backfills, unattended agents
Where that work lands
85%
Of the deferrable spend. Six venues publish this tier
0%
Regional and time-bound promos. Remainder stays at spot

You would stop paying

$6,800

per month · 17.0% off the whole bill

Per year$81,600
Deferrable spend$16,000 / mo
New monthly bill$33,200 / mo
Discount on that work42.5%

Every number here is arithmetic on a published price sheet — no modeled inputs, no assumed energy, nothing that needs a footnote. Actual capture depends on how much of your work really has slack.

02 // The sheet

What the prices are

offpeak.prices // snapshot 2026-08-30 · USD per 1M tokens

ModelStandard inStandard outBatch / off-peak inBatch / off-peak outFast inFast out

gpt-5.6-sol's standard rate is promotional at least through 2026-11-21; post-promo list is $5 / $30 and both tiers move with it, so the ratio is the durable figure, not the dollars. Batch is exactly 50% of standard. The openai/gpt-oss rows are Groq's, the mistral rows are from mistral.ai/pricing/api, and the gemini rows from ai.google.dev/pricing — all three at the same 50% batch rule. Groq and Mistral publish no per-model fast rate, so this sheet carries none; Google prices a priority tier that this sheet does not carry yet, and tools/sheet_reconcile.py reports the gap rather than the sheet implying a spread it has not verified. Gemini 3.7 and 3.6 Flash are introductory through 2026-12-31 and then double to $1.50 / $7.50 on both legs, which moves the dollars and leaves the batch ratio alone. The deepseek rows (api-docs.deepseek.com) carry the peak rate as standard; DeepSeek has no batch API, and its half price is the off-peak clock — peak is 01:00–04:00 and 06:00–10:00 UTC on weekdays, everything else and all weekend is half — so the discount is decided per request by the wall clock, not by a queue. The qwen rows are Alibaba Model Studio's international region, batch at 50%; qwen3.7-max is "limited-time 50% off" with no published end date, so it carries no expiry and may step up unannounced. Fast tiers are published on the gpt-5.6 family and on Claude Opus 5 / Opus 4.8 ($10 / $50, research preview, first-party API only, not combinable with batch) — fastest to batch is at both venues, computed live by urgency_spread(), never cached. Sonnet 5's $2 / $10 is now Anthropic's standard price; the scheduled 09-01 rise to $3 / $15 will not occur.

03 // Honest caveats

What this does not prove

ESTIMATE

Not a quote

Actual capture depends on what share of your work really has slack, and on rates holding. A quote comes from your own metadata, free, in the flexibility report.

RESCUES

Insurance costs

A job rescued at its deadline pays list, not batch. Smoke runs to date: zero fallbacks, 6/6 SLAs — but the model above assumes none, and real books have some.

Quoting daily

Spread Board

Offpeak's public track record

What waiting is worth, in public: the three spreads a deadline can capture — tokens, power, carbon — quoted from open data at 19:00Z, marked against actuals at 06:30Z, daily, against the deadline and not the clock.

2.0×
Tokens
batch at half of list
1.38×
Power · GB
day-ahead, this session
1.44×
Carbon · GB
chosen window

QUOTES COST NOTHING · SETTLEMENTS ONLY EVER COME FROM REAL RUNS

01 // The live quote

Three legs, latest session

Tokens // both venues

2.0×

Batch at 50% of list on both published sheets. Fast → batch, at both venues that sell one: .

Power // GB · Agile day-ahead

1.38×

34.85p across the 17–21 peak → 25.23p in the 00–05 window. Extremes 40.99p / 23.61p.

Carbon // GB · NESO forecast

1.44×

Cleanest 5h 110.9g · dirtiest 159.5g. The naive 00–05 shift is 0.74× — worse than peak.

Full brief
WORKED: A 12.68M-TOKEN CORPUS — $21.12 RUN NOW, $10.56 ON A DEADLINE. SAME MODELS, SAME TOKENS. THE FASTEST→BATCH RANGE ON EITHER VENUE THAT SELLS A FAST TIER IS CURRENTLY 4×, PART OF IT PROMOTIONAL — RE-DERIVED EACH SESSION, NEVER CACHED. POWER IS KNOWN, NOT FORECAST — DAY-AHEAD AUCTIONS CLEAR BEFORE DELIVERY. ONLY THE CARBON LEG IS A FORECAST AT QUOTE TIME, AND IT IS MARKED AGAINST ACTUALS AT 06:30Z. THIS SESSION THE CARBON LEG IS INVERTED ON THE NAIVE WINDOW AND IS PUBLISHED THAT WAY.

02 // Marked

Marked against actuals

Marked sessions // actuals

Leg · zoneHigh · latestLow · latestPrior sessionLatest markedNote
POWER · GBAGILE, DAY-AHEAD37.17p24.82p1.43×1.50×extremes 1.74× and 2.28×
POWER · CAISO SP15DA LMP $/MWh$82.52$50.061.27×1.65×evening peak, solar rolled off
POWER · ERCOT HOUDA LMP $/MWh$93.83$23.703.94×3.96×widest zone on the board, twice
CARBON · GBNESO · CHOSEN WINDOW146.1 g65.0 g1.13×2.25×naive 00–05 reads 1.03× and 2.10×
CARBON · CISO + ERCOEIA-930 · DERIVED · BACKFILLED0.80× / 0.80×both sub-parity; the prior session awaits its following-morning trough
TOKENS · OAI + ANTPUBLISHED SHEETSlist50% list2.0×2.0×constant until a vendor moves a sheet — which this board catches
CAISO · MARKEDCARBON SPREAD 0.76× — BELOW 1. Evening solar is gone by midnight: cheap hours are not automatically clean hours. Every carbon claim here is computed per-hour.
Methodology & sources
SOURCES, OPEN AND KEYLESS WHERE NOTED: GB CARBON — NESO CARBON INTENSITY API (KEYLESS, WITH FORECASTS) · GB POWER — OCTOPUS AGILE DAY-AHEAD (KEYLESS, REGION C) · TOKEN PRICES — OPENAI + ANTHROPIC PUBLISHED SHEETS, SNAPSHOT 2026-08, ENCODED IN offpeak.prices · US ZONES — CAISO + ERCOT DA LMP VIA gridstatus · US CARBON — EIA-930 HOURLY GENERATION MIX, KEY WIRED 08-22; A SESSION'S SPREAD IS MEASURED AGAINST THE FOLLOWING MORNING'S TROUGH, SO THE DATA LANDS ~TWO DAYS AFTER THE MARK AND EVERY MARK BACKFILLS THE SESSIONS THAT WERE TOO EARLY — UNTIL THEN THEY READ UNAVAILABLE-AND-DOCUMENTED. A DAILY JOB QUOTES 19:00Z AND MARKS 06:30Z, COMMITTING PER-SESSION JSON TO THE OPEN REPO — EVERY NUMBER RECOMPUTABLE FROM CITED PUBLIC ENDPOINTS: github.com/offpeak-ai/offpeak

03 // The carbon leg

Cheap hours are not clean hours

Exhibit 03 // GB carbon · every session marked

The board's third leg is the one that misbehaves. Run the naive version — "shift it to 00:00–05:00" — and the carbon result is a coin flip: 1.03×, 2.10×, and 0.74×. That last one means the window was dirtier than the evening peak. Evening solar is gone by midnight and what backfills it is not clean. CAISO has done the same, at 0.76×.

Pick the window from the curve instead of assuming it and the same sessions read 1.13×, 2.25× and 1.44× — never below parity. Price and carbon are two different curves and get optimized separately.

WHAT WE DO NOT CLAIMOn an API batch tier, Offpeak does not choose the hour or the region — the provider does. So this is a measurement of what a scheduler with that control could capture, not a carbon saving we book on your behalf. That claim waits for the metered self-hosted lane.
Definitions & sources
NAIVE = MEAN gCO2/kWh OVER 17:00–21:00 LOCAL ÷ MEAN OVER 00:00–05:00 — THE RATIO A CRON GETS. CHOSEN = DIRTIEST CONTIGUOUS 5-HOUR MEAN ÷ CLEANEST, INSIDE THE SAME 16:00Z–07:00Z SESSION — THE RATIO A SCHEDULER GETS WHEN IT READS THE CURVE. BOTH COMPUTED BY tools/night_report.py FROM ONE HALF-HOURLY SERIES AND COMMITTED AS PER-SESSION JSON. GB CARBON: NESO CARBON INTENSITY API, KEYLESS. US CARBON: EIA-930, KEY WIRED 08-22. A SESSION'S SPREAD NEEDS THE FOLLOWING MORNING'S 00:00–05:00 LOCAL HOURS, SO THE FEED IS ROUGHLY TWO DAYS BEHIND THE MARK, NOT ONE — EVERY MARK NOW GOES BACK AND BACKFILLS THE SESSIONS THAT WERE TOO EARLY. 08-20 HAS FILLED: CISO 0.80×, ERCO 0.80×. 08-21 AND 08-22 STILL RECORD UNAVAILABLE-AND-DOCUMENTED, WITH THE DAY THEY ARE WAITING ON NAMED, RATHER THAN A NUMBER NOBODY MEASURED. n = 2 MARKED SESSIONS + 1 QUOTED. THE SUB-PARITY RESULT IS PUBLISHED THE SAME WAY THE GOOD ONES ARE.
Evidence 03 · carbon spread · multiple
Naive 00–05 Chosen window Worse than peak

Anything left of parity is a session where shifting the work would have added carbon. The chosen window has never landed there.

NESO · keyless · marked 06:30Z

04 // Settlements

Real runs only

The settled record // real money, real venues · SETTLED.md

RunVenuesJobsListPaidCapturedSLA
mechanics-1TWO VENUES · OPENAI LEG REJECTEDanthropic + openai batch48$0.00156$0.00078150.0%24/48
mechanics-2CEILING TOO LOW · EMPTY OUTPUTSopenai batch24$0.000594$0.00029750.0%24/24
mechanics-3CEILING SIZED TO THE MODELopenai batch24$0.00131$0.00065550.0%24/24
groq-1BATCH TIER GATED · SYNC FALLBACK AT LISTgroq · sync fallback24$0.00236$0.002360.0%24/24
mistral-1BATCH TIER GATED · SYNC FALLBACK AT LISTmistral · sync fallback24$0.000199$0.0001990.0%24/24
gemini-1BATCH TIER · FIRST THIRD VENUE TO CAPTUREgemini batch5$0.00919$0.0046050.0%5/5
08-26 openai-1BATCH TIERopenai batch24$0.00127$0.00063650.0%24/24
08-26 anthropic-1BATCH TIERanthropic batch6$0.000378$0.00018950.0%6/6
08-26 gemini-1BATCH TIER · REPEATED TWO DAYS ONgemini batch5$0.00918$0.0045950.0%5/5
08-26 mistral-2BATCH COMPLETED, THEN DISCARDED · CLIENT BUGmistral · sync fallback24$0.000198$0.0001980.0%24/24
08-26 mistral-3BATCH TIER · FOURTH VENUE TO CAPTUREmistral batch24$0.000200$0.000099950.0%24/24

Scale is on every row on purpose: these are mechanics proofs, not production volume — the evidence they carry is that the submit → poll → settle path works and prices out at exactly half of list on a real bill. Three rows captured nothing, and they stay. Groq (403 not_available_for_plan) still gates the batch API behind a plan, so those jobs took the sync fallback and paid list. Mistral was gated the same way (402) until billing was switched on — and then failed a second time for a different reason: on 08-26 the batch completed and was thrown away by a bug in offpeak's own driver, which reached for a streaming response's text before reading it. That run paid twice for the same answers and captured nothing. It is on the board next to the fixed re-run, because a client that loses a batch costs exactly what a venue that refuses one costs, and is harder to notice. The full record.

Real money, published whole

Settled runs

The settled record

Real runs on real bills, across five venues — four of them now capturing: OpenAI, Anthropic, Google Gemini and Mistral. Every run that reached a batch tier and kept it captured exactly 50.0% against the published sheet. Runs that captured nothing are published anyway — a gated venue, and one night where the batch completed and offpeak's own driver threw it away. The failures are the reason to believe the rest. Scale is printed on every row on purpose: this ledger proves the path, not the volume.

50.0%
Captured on every run
that kept its batch
4
Venues capturing
at the batch tier
100%
Of failures published
rather than dropped

01 // The record

The record, whole

Receipts // generated by tools/mechanics_run.py + settle_report.py · SETTLED.md

RunVenuesJobsTokensListPaidCapturedSLA
mechanics-148 JOBS, TWO VENUESanthropic 24 · openai 2448787 / 155$0.00156$0.00078150.0%24/48
mechanics-2CEILING TOO LOWopenai batch24728 / 374$0.000594$0.00029750.0%24/24
mechanics-3CEILING SIZED TO THE MODELopenai batch24728 / 971$0.00131$0.00065550.0%24/24
groq-1BATCH TIER GATED · SYNC FALLBACK AT LISTgroq · sync fallback242,288 / 7,307$0.00236$0.002360.0%24/24
mistral-1BATCH TIER GATED · SYNC FALLBACK AT LISTmistral · sync fallback24948 / 94$0.000199$0.0001990.0%24/24
gemini-1BATCH TIER · FIRST THIRD VENUE TO CAPTUREgemini batch5124 / 2,427$0.00919$0.0046050.0%5/5
08-26 openai-1BATCH TIER · 8m07sopenai batch24728 / 938$0.00127$0.00063650.0%24/24
08-26 anthropic-1BATCH TIER · 4m05santhropic batch6188 / 38$0.000378$0.00018950.0%6/6
08-26 gemini-1BATCH TIER · 2m04sgemini batch5124 / 2,422$0.00918$0.0045950.0%5/5
08-26 mistral-2BATCH COMPLETED, THEN DISCARDED · CLIENT BUGmistral · sync fallback24948 / 93$0.000198$0.0001980.0%24/24
08-26 mistral-3BATCH TIER · FOURTH VENUE TO CAPTURE · 33smistral batch24948 / 96$0.000200$0.000099950.0%24/24

Tokens shown as in / out. Every figure is emitted by the settlement code from counted tokens against the published sheet — none of it is typed in by hand, and the arithmetic reproduces in the calculator.

02 // What went wrong

Five failures, published

RUN 1 · HALF REJECTED

The venue said no

Anthropic settled 24/24. OpenAI rejected all 24 with an HTTP 400 — the driver sent a token-ceiling parameter the gpt-5.6 family no longer accepts. Nothing was billed on the failed leg.

SLA 24/48 · $0 on the rejected venue
RUN 2 · EMPTY ANSWERS

Billed for nothing useful

The OpenAI leg re-ran and settled 24/24 — and returned 24 empty strings, because a 16-token ceiling was spent entirely on reasoning. Billed, SLA met, output worthless. Published rather than dropped.

SLA 24/24 · 374 output tokens · 0 answers
RUN 3 · CLEAN

Same lines, right ceiling

The same 24 lines at a 256-token ceiling: 24/24 real answers, 971 output tokens, zero fallbacks, every deadline met, half of list paid.

SLA 24/24 · captured $0.000655 · 50.0%
RUN 4 · TIER NOT ON SALE

The venue was gated

Groq's entire batch surface answered 403 not_available_for_plan — file upload and batch listing alike — so submission never completed and all 24 jobs took the sync fallback at list price. Billed, every deadline met, nothing captured. The blocker is plan entitlement, not the driver.

SLA 24/24 · $0.00236 paid · 0.0% captured
RUN 5 · GATED AGAIN

And the probe missed it

Mistral answered 402 · enable billing via the console on batch create. Worse, an entitlement probe had cleared it a day earlier: a malformed request came back with a field-level validation error, so it looked reachable. Validation runs before the billing check — a request that cannot succeed can never reach the paywall.

SLA 24/24 · $0.000199 paid · 0.0% captured
RUN 6 · A THIRD VENUE

Gemini captures

Google Gemini batch, once billing was enabled: 5/5, zero sync fallbacks, 50.0% captured, 2m36s end to end. The first venue here beyond OpenAI and Anthropic to collect the spread rather than record why it could not.

SLA 5/5 · captured $0.00460 · 50.0%
RUN 7 · WE LOST IT

The batch worked. We didn't.

Mistral's paywall was gone and the batch succeeded — 24/24 in 6m16s. offpeak threw it away: _download reached for a streaming response's .text before reading it, which raises ResponseNotRead — a RuntimeError, so the getattr(…, None) guard never applied. run() booked a finished batch as a polling failure and paid list for all 24. It had completed 32 seconds earlier.

SLA 24/24 · paid twice · 0.0% captured
RUN 8 · A FOURTH VENUE

Mistral captures

Same twenty-four lines, ninety minutes later, on the fixed driver: 24/24, zero fallbacks, 50.0% captured, 33 seconds end to end. Nothing about the venue changed — only whether offpeak could read the file it was already being handed.

SLA 24/24 · captured $0.0000999 · 50.0%
LESSONThe sync fallback rescues jobs that never came back — not jobs that came back failed. Run 1 found that distinction the only way it can be found, and the driver now carries the fix. Run 7 found the sharper version: a fallback that fires on a batch which already succeeded is not a rescue, it is a second bill. Both were found by spending real money and publishing what it bought.

03 // The parked book

The 6,000-job exhibit

Re-priced against current published rates, the full demonstration book came to $28.05 list against a $23.24 authorized ceiling — over, so it was never built and never submitted. It stays parked until a partner conversation needs the exhibit, or until a partner's own runner settles it on their keys, which costs Offpeak nothing.

Job classModel · venueJobsTokens / jobList $Batch $
Summarize 5k-token doc chunksANTHROPIC BATCHclaude-sonnet-51,0005,000 / 30013.006.50
Summarize 5k-token doc chunksOPENAI BATCHgpt-5.6-terra1,0005,000 / 30013.606.80
Classify short passagesANTHROPIC BATCHclaude-haiku-4-52,000500 / 201.200.60
Classify short passagesOPENAI BATCHgpt-5.6-luna2,000500 / 200.250.12
Re-priced at current rates~12.7M TOKENS6,00028.0514.02

03 // Honest caveats

What this record does not prove

SCALE

Sub-cent is sub-cent

The captured dollars are deliberately tiny. What is proved is the path and the arithmetic, on real bills. Production volume is a different claim and is not made here.

OWN KEYS

Offpeak's own work

These ran on Offpeak's keys against Offpeak's own book. No buyer receipt is on this ledger yet — that is the open half of the proof, and the point of the design-partner program.

RESCUES

Insurance costs

A job rescued at its deadline pays list, not batch. Zero fallbacks on the clean runs — but run 1 showed the fallback has a blind spot, and a real book will find more.

Sheets watched daily

Prices

Every number the desk quotes against, drawn

Three series, one page: the token sheet every venue publishes and the desk prices against, the power marks that show the same valley in three grids, and the batch turnaround our canaries measure — the number no sheet prints.

4.0×
Fast → batch
same model, same tokens
Power · ERCOT
peak over off-peak, latest marked
Batch turnaround
median, completed canaries

Sheet snapshot · board-data as committed

01 // The sheet

Same model, three prices

Every row is one model at one venue. The dim end is the batch tier — or, on DeepSeek, the off-peak clock rate, half of peak by time of day rather than by queue — the bright end is standard, and where a venue sells a fast tier it sits further right. The gap is what waiting is worth — log scale, because the sheet spans three orders of magnitude and the ratio is the durable figure, not the dollars.

Exhibit 01 · price sheet · $/M tokens

Batch → standard → fastoutput · $/M tokens

Batch · half of standard Standard Fast · gpt-5.6 family · Opus 5 / 4.8 (preview) Off-peak · DeepSeek's clock lane · half of peak

SOURCE: offpeak.prices, PROVIDER PUBLIC SHEETS, SNAPSHOT 2026-08-30, RE-VERIFIED FROM RENDERED PAGES. BATCH = EXACTLY HALF OF STANDARD ON EVERY ROW SHOWN. DEEPSEEK SELLS NO BATCH: ITS STANDARD ROW IS THE PEAK RATE (01:00–04:00 + 06:00–10:00 UTC, MON–FRI) AND THE DIM END IS THE OFF-PEAK RATE, HALF OF PEAK BY THE WALL CLOCK — SAME ARITHMETIC, DIFFERENT LANE. QWEN (ALIBABA MODEL STUDIO, INTERNATIONAL REGION) BATCHES AT 50%; QWEN3.7-MAX IS "LIMITED-TIME 50% OFF" WITH NO PUBLISHED END DATE. FAST = 2× STANDARD WHERE SOLD: THE GPT-5.6 FAMILY, AND CLAUDE OPUS 5 / OPUS 4.8 AS A RESEARCH PREVIEW (FIRST-PARTY API ONLY, NOT WITH BATCH) — 4× FAST → BATCH AT BOTH VENUES. SHEET RE-VERIFIED AGAINST THE RENDERED PROVIDER PAGES 2026-08-30. PROMOTIONAL RATES (SOL ≥ 2026-11-21 · GEMINI 3.6/3.7 FLASH ≤ 2026-12-31) CARRY THEIR EXPIRY AND RE-DERIVE, NEVER CACHED. CACHED-INPUT AND LONG-CONTEXT COLUMNS ARE NOT DRAWN.

02 // The marks

The same valley, three grids

Power is the leg with a public day-ahead price. Each session the board quotes the 17:00–21:00 local evening peak against the 00:00–05:00 trough and marks it the next morning. The ratio panel puts all three grids on one axis; the level panels show the dollars behind it, in each grid's own unit.

Exhibit 02 · window spread · marked sessions

Peak over off-peakratio · marked at 06:30Z

GB · Octopus Agile · region C

GBp/kWh

CAISO · SP15 · day-ahead LMP

CAISO$/MWh

ERCOT · Houston · day-ahead LMP

ERCOT$/MWh

SOURCES: GB — OCTOPUS AGILE DAY-AHEAD (KEYLESS, REGION C), HALF-HOURLY. CAISO SP15 AND ERCOT HOUSTON — DAY-AHEAD LMP VIA gridstatus, HOURLY. A SESSION SPANS 16:00Z–07:00Z; THE OFF-PEAK WINDOW LANDS THE FOLLOWING LOCAL MORNING, SO THE US LEGS MARK ONE DAY BEHIND GB. GAPS ARE SESSIONS WHOSE TROUGH HOURS HAD NOT LANDED WHEN THE MARK WAS TAKEN — RECORDED AS UNAVAILABLE, NEVER FILLED IN. EACH POINT IS ONE nightly/<session>-mark.json ON BOARD-DATA; RELATIVE LABELS ON THE AXIS, THE SESSION DATE ON HOVER AND IN THE TABLE.

03 // The queue

How long the cheap lane takes

No venue publishes this. Every session the desk submits a two-job canary into each venue's batch tier with a 24-hour window and records how long it actually took. Most clear in minutes; one took eight and a half hours. A hollow marker at the top is a censored leg — still running when the session closed, so all we know is "longer than we watched."

Coverage // what is measured, and what is not · as of 2026-08-30

The chart below draws only venues with a measured queue — four today. DeepSeek's clock lane is measured differently: there is no batch to wait for, so the probe records which regime the clock is in and the synchronous latency of a request made in it. Hourly rows stay in the private probe repo; this page carries the daily series, per-venue percentiles and counts — never hour-of-day.

VenueLaneDiscountStatusSinceNote
OpenAIBATCH APIbatch50% offMEASUREDdaily 08-23 · hourly 08-30Capturing. Two-job canary, 24 h window; the daily series is public, the hourly rows are private.
AnthropicMESSAGE BATCHESbatch50% offMEASUREDdaily 08-23 · hourly 08-30Capturing. Same canary, same window.
Google GeminiBATCH MODEbatch50% offMEASUREDdaily 08-26 · hourly 08-30Capturing. Same canary, same window.
MistralBATCH APIbatch50% offMEASUREDdaily 08-26 · hourly 08-30Capturing. Same canary; the one eight-and-a-half-hour turnaround on the chart is this venue's.
DeepSeekCLOCK LANE · NO BATCH APIclock50% off-peakMEASUREDhourly 08-30Not a queue, so not on the chart. Observable is regime plus synchronous latency: first two observations 2.3 s and 2.6 s, off-peak, 2,000 / 300 tokens.
QwenALIBABA MODEL STUDIO · INTLbatch50% offPROBATIONfirst live run 08-30Driver shipped in 0.2.7. Joins the chart once it has settled repeatedly.
GroqBATCH APIbatch50% offBLOCKEDBatch tier answers 403 — plan-gated, and plan sales are suspended. Sync fallback only.
xAIFLAGSHIP−20% legacy onlyNO LANENo batch lane on the flagship; the discount is on legacy models only. Watched, not probed.
AWS BedrockBATCH INFERENCEbatch50% offNOT BUILTPublished; driver on partner pull.
AzureBATCHbatchNOT BUILTOn partner pull.
KimiMOONSHOTbatch40% offNOT BUILTPublished at 40%; not built.
Exhibit 03 · batch turnaround · canaries

Turnaround by venuelog scale · 24 h window

Inline series · the floor

SOURCE: THE DAILY SERIES ABOVE IS nightly/<session>-queue.json ON BOARD-DATA AS COMMITTED, WRITTEN BY tools/queue_probe.py; THE PROBE NOW RUNS HOURLY FROM A PRIVATE REPO AND PUBLISHES BACK nightly/QUEUE.md AND nightly/queue-summary.json — PER-VENUE PERCENTILES AND COUNTS, NEVER HOUR-OF-DAY, NEVER A PER-ROW HOURLY SERIES. TWO JOBS PER LEG, ~30 TOKENS IN, SUB-CENT SPEND CAPPED AT $0.01 PER SESSION. SKIPPED LEGS (NO KEY IN THE ENVIRONMENT) ARE NOT DRAWN AND STAY IN THE TABLE. THIS IS THE REALIZED HALF OF THE RECORD — THE SHEET SAYS "WITHIN 24 HOURS"; THIS SAYS WHAT THAT MEANT.

Link copied