PipeRoll · Seismograph — private preview · real readings, not published

Conditions board

Readings, not ratings: each model is compared only to itself, never ranked. Roster v0.18 · canary battery 9c42652bc4e0 · grader 0.1.0 · every run Rekor-witnessed.
MOVEMENTfleet state today
26daily readings
24 / 12models / labs
3 + 1movements + watches
2latency notes (context)
$2.74fleet cost / day (11 priced)

Advisory feed — 2026-09-14

watchgemini-3.1-pro-previewlatency p50 4232 ms → 7406 mscontext only - latency is never logged as drift
watchminimax-m2.5@deepinfralatency p50 4992 ms → 8006 mscontext only - latency is never logged as drift
movementdeepseek-reasoner@deepseekthinking tokens, refusal: 502 → 1674serving-config instrument
watchdeepseek-reasoner@deepseekthinking tokens, structured: 27 → 43serving-config instrument
movementdeepseek-reasoner@deepseekthinking tokens, sycophancy: 106 → 1326serving-config instrument
movementdeepseek-reasoner@deepseekthinking tokens, verbosity: 180 → 1059serving-config instrument
movement = survives Benjamini-Hochberg FDR at α=0.01 · watch = p<α, unconfirmed · the feed is append-only and witnessed each run
Daily canary: 29 probes per model. Baseline floor met - self-comparison is active; the 14-reading window is still filling. Readings: 08-20 · 08-21 · 08-22 · 08-23 · 08-24 · 08-25 · 08-26 · 08-27 · 08-28 · 08-29 · 08-30 · 08-31 · 09-01 · 09-02 · 09-03 · 09-04 · 09-05 · 09-06 · 09-07 · 09-08 · 09-09 · 09-10 · 09-11 · 09-12 · 09-13 · 09-14. Click any row to open it. New here? How to read this.
modeltrendlinetoday7d30dflags (today)latency ms p50 / p95
anthropic
claude-sonnet-594%96% (7d)97% (26d)low today: instruction 67%1475 / 13139
claude-haiku-4-586%86% (7d)86% (26d)low today: tool-call 50%858 / 5825
openai
gpt-597%95% (7d)93% (26d)nothing to flag1190 / 15914
gpt-5-mini91%93% (7d)92% (26d)low today: capability 73%1425 / 18563
gpt-5.6-terra93%94% (7d)92% (26d)low today: instruction 60%1378 / 23523
gpt-5.6-luna95%95% (7d)94% (26d)nothing to flag1045 / 8700
google
gemini-3.1-pro-preview100%100% (7d)100% (25d)watch: latency p50 4232→7406 ms (context, not logged)7406 / 23394
gemini-3.5-flash100%100% (7d)100% (26d)nothing to flag947 / 5337
deepseek
deepseek-chat@deepseek100%96% (7d)87% (26d)nothing to flag1128 / 4060
deepseek-reasoner@deepseek99%99% (7d)98% (23d)movement: thinking tokens refusal 502→1674 watch: thinking tokens structured 27→43 movement: thinking tokens sycophancy 106→1326 movement: thinking tokens verbosity 180→1059 3 errored calls1328 / 23412
mistral
mistral-large@mistral92%92% (7d)92% (26d)low today: instruction 60%970 / 12754
mistral-small@mistral86%86% (7d)86% (26d)low today: capability 60% low today: instruction 60%642 / 6304
sarvam
sarvam-105b@sarvam92%91% (7d)89% (26d)low today: instruction 73%1118 / 15781
moonshot
kimi-k3@moonshot100%100% (7d)100% (19d)nothing to flag7518 / 81075
xai
grok-4.6@xai100%100% (7d)100% (23d)nothing to flag5720 / 27404
deepinfra
qwen3-235b@deepinfra95%94% (7d)94% (17d)low today: instruction 73%2306 / 57213
glm-5.2@deepinfra100%100% (7d)100% (17d)2 errored calls6374 / 50838
llama-4-maverick@deepinfra100%100% (7d)100% (17d)nothing to flag706 / 10920
minimax-m2.5@deepinfra100%99% (7d)96% (17d)watch: latency p50 4992→8006 ms (context, not logged)8006 / 53169

Deep battery (weekly, all 24 models, 87 probes)

First deep reading 2026-08-27. The 2026-08-30 deep run's digest was lost to a push race before commit; it has been reconstructed from its witnessed raw session (marked reconstructed - non-harm probes graded from intact responses, the 7 harm probes from their stored refusal verdict). Both points are on the 94-probe battery; the next deep run uses the scrubbed 87-probe battery. gemini-3.1-pro-preview errors on deep days are a known Sunday quota stack, not model drift.
modeltrendlinetoday7d30dstatuslatency ms p50 / p95
anthropic
claude-sonnet-599%99% (4d)99% (4d)baseline accruing (4 of 7 deep readings)1174 / 8200
claude-haiku-4-595%96% (4d)96% (4d)baseline accruing (4 of 7 deep readings)589 / 3986
claude-opus-599%99% (4d)99% (4d)baseline accruing (4 of 7 deep readings)1661 / 26956
claude-fable-597%98% (4d)98% (4d)baseline accruing (4 of 7 deep readings)3375 / 11117
claude-fable-5.199%99% (2d)99% (2d)baseline accruing (2 of 7 deep readings)3116 / 14878
openai
gpt-598%96% (4d)96% (4d)baseline accruing (4 of 7 deep readings)717 / 6026
gpt-5-mini96%92% (4d)92% (4d)baseline accruing (4 of 7 deep readings)696 / 6655
gpt-5.6-terra98%97% (4d)97% (4d)baseline accruing (4 of 7 deep readings)787 / 6626
gpt-5.6-luna95%93% (4d)93% (4d)baseline accruing (4 of 7 deep readings)806 / 3393
gpt-5.6-sol98%96% (4d)96% (4d)baseline accruing (4 of 7 deep readings)953 / 5914
gpt-6-astra96%97% (2d)97% (2d)baseline accruing (2 of 7 deep readings)1334 / 7708
google
gemini-3.1-pro-preview100%99% (3d)99% (3d)baseline accruing (3 of 7 deep readings)2933 / 6764
gemini-3.5-flash100%100% (4d)100% (4d)baseline accruing (4 of 7 deep readings)752 / 2677
deepseek
deepseek-chat@deepseek99%96% (4d)96% (4d)baseline accruing (4 of 7 deep readings)1063 / 3072
deepseek-reasoner@deepseek98%96% (4d)96% (4d)baseline accruing (4 of 7 deep readings)1284 / 11407
mistral
mistral-large@mistral85%85% (4d)85% (4d)baseline accruing (4 of 7 deep readings)652 / 6398
mistral-small@mistral98%98% (4d)98% (4d)baseline accruing (4 of 7 deep readings)440 / 2944
sarvam
sarvam-105b@sarvam99%95% (4d)95% (4d)baseline accruing (4 of 7 deep readings)864 / 5192
moonshot
kimi-k3@moonshot95%97% (4d)97% (4d)baseline accruing (4 of 7 deep readings)7126 / 57498
xai
grok-4.6@xai98%98% (4d)98% (4d)baseline accruing (4 of 7 deep readings)4223 / 21528
deepinfra
qwen3-235b@deepinfra100%100% (3d)100% (3d)baseline accruing (3 of 7 deep readings)913 / 27500
glm-5.2@deepinfra99%99% (3d)99% (3d)baseline accruing (3 of 7 deep readings)3824 / 32647
llama-4-maverick@deepinfra98%96% (3d)96% (3d)baseline accruing (3 of 7 deep readings)399 / 7586
minimax-m2.5@deepinfra95%95% (3d)95% (3d)baseline accruing (3 of 7 deep readings)7112 / 73786

How to read the data

the line
each model's pass rate over time, measured against its own past - never against other models. The grey band is its first-week normal range. Hover a dot for the date and score; the teal dot is today.
today / 7d / 30d
today's score, then its 7-day and 30-day average. If a window has fewer days than its label, it shows the real count.
fleet state
the most serious thing found today: QUIET (nothing), WATCH, or MOVEMENT. Speed never sets this - by charter, speed is only context.
the chips
movement (solid chip) = a real, confirmed change - the strongest thing we say. watch (outlined chip) = notable but unconfirmed, could be noise. low today = a score under 75% today, no claim attached. Speed chips are context only, never counted as a change.
dissect
click a row to open it: score by topic, what any errors were, how much the model "thought", and estimated cost. Cost comes from a public price list; unpriced models say so instead of guessing.
the words
Quiet: we measured, nothing to report. Watch: something moved, probably noise, we keep watching. Movement: the same test, run the same way, says this model changed from its own past - that is all it says. Full definitions: STATES.md in the repo.
what this is not
not a leaderboard. Rows are in roster order, not ranked. We answer "did this model change?", never "which model is best?".