SmallDecide v0.3.0 ·
Evaluation results
SmallDecide is a 0.6B model that scores a supplied set of answers. On a 500-query routing test, it answered 486 correctly in 161 ms median on an M4 Max. The comparisons below cover 33 hosted models, a confidence gate, and a Factorio integration.
Routing results
Each query has eight candidate intents. SmallDecide scores 97.2%, with a 161 ms median and 222 ms p95. Allocating $0.50/hour to a fully utilized M4 Max gives $0.02307 per 1,000 decisions.
Fable 5, GPT-6 Astra, and DeepSeek V4 Pro score higher, at 99.4–99.8%, with longer API response times. Across the full panel, Llama 3.1 8B is fastest at about 100 ms (96.0% accuracy); Command R7B is cheapest at $0.00482 per 1,000 (97.0%).
Full results: 33 hosted models + SmallDecide
| Model / run | Correct / 500Accuracy · 95% Wilson interval | Medianms | p95ms | Cost / 1,000USD |
|---|---|---|---|---|
| SmallDecide 0.6BM4 Max · warm, sequential | 486/50097.2% · 95.36%–98.32% | 161 | 222 | $0.02307 |
| Amazon: Nova Lite 1.0Original panel | 493/50098.6% · 97.14%–99.32% | 476 | 2,969 | $0.00860 |
| Amazon: Nova Micro 1.0Original panel | 491/50098.2% · 96.61%–99.05% | 409 | 545 | $0.00502 |
| Anthropic: Claude Haiku 4.5Original panel | 494/50098.8% · 97.41%–99.45% | 778 | 981 | $0.16044 |
| Anthropic: Claude Sonnet 5Original panel | 493/50098.6% · 97.14%–99.32% | 1,719 | 2,111 | $0.38084 |
| Claude Fable 5Leading-model run · low reasoning | 497/50099.4% · 98.25%–99.80% | 3,209 | 4,232 | $1.93940 |
| Cohere: Command R7B (12-2024)Original panel | 485/50097.0% · 95.11%–98.17% | 228 | 319 | $0.00482 |
| DeepSeek V4 Pro (0813)Leading-model run · low reasoning | 499/50099.8% · 98.88%–99.96% | 3,294 | 12,529 | $0.44780 |
| DeepSeek: DeepSeek V4 Flash 0731Original panel | 491/50098.2% · 96.61%–99.05% | 567 | 5,948 | $0.00802 |
| Google: Gemini 2.5 FlashOriginal panel | 493/50098.6% · 97.14%–99.32% | 468 | 662 | $0.04047 |
| Google: Gemini 2.5 Flash LiteOriginal panel | 494/50098.8% · 97.41%–99.45% | 424 | 558 | $0.01301 |
| Google: Gemini 3.1 Flash LiteOriginal panel | 496/50099.2% · 97.96%–99.69% | 656 | 922 | $0.03277 |
| Google: Gemma 3 27BOriginal panel | 496/50099.2% · 97.96%–99.69% | 292 | 598 | $0.01417 |
| GPT-6 AstraLeading-model run · low reasoning | 498/50099.6% · 98.55%–99.89% | 1,349 | 2,721 | $1.58390 |
| Meta: Llama 3.1 8B InstructOriginal panel | 480/50096.0% · 93.90%–97.40% | 100 | 158 | $0.03069 |
| Meta: Llama 3.2 1B InstructOriginal panel | 49/5009.8% · 7.49%–12.72% | 2,342 | 9,875 | $0.00858 |
| Meta: Llama 3.2 3B InstructOriginal panel | 439/50087.8% · 84.64%–90.38% | 224 | 391 | $0.00877 |
| Meta: Llama 3.3 70B InstructOriginal panel | 494/50098.8% · 97.41%–99.45% | 416 | 1,669 | $0.02972 |
| Meta: Llama 4 MaverickOriginal panel | 494/50098.8% · 97.41%–99.45% | 575 | 932 | $0.02669 |
| Mistral: Ministral 3 3B 2512Original panel | 457/50091.4% · 88.62%–93.55% | 373 | 695 | $0.01212 |
| Mistral: Mistral Large 3 2512Original panel | 497/50099.4% · 98.25%–99.80% | 628 | 1,161 | $0.06944 |
| Mistral: Mistral NemoOriginal panel | 483/50096.6% · 94.62%–97.87% | 312 | 551 | $0.00551 |
| Mistral: Mistral Small 3.2 24BOriginal panel | 477/50095.4% · 93.19%–96.92% | 700 | 2,943 | —¹ |
| NVIDIA: Nemotron 3.5 LightningOriginal panel | 480/50096.0% · 93.90%–97.40% | 137 | 226 | $0.01051 |
| OpenAI: GPT-4.1Original panel | 495/50099.0% · 97.68%–99.57% | 656 | 1,028 | $0.28076 |
| OpenAI: GPT-4.1 MiniOriginal panel | 496/50099.2% · 97.96%–99.69% | 667 | 1,065 | $0.05615 |
| OpenAI: GPT-4.1 NanoOriginal panel | 494/50098.8% · 97.41%–99.45% | 605 | 1,011 | $0.01404 |
| OpenAI: GPT-4o-miniOriginal panel | 492/50098.4% · 96.87%–99.19% | 502 | 685 | $0.02048 |
| OpenAI: GPT-5.6 LunaOriginal panel | 491/50098.2% · 96.61%–99.05% | 1,145 | 3,423 | $0.03228 |
| OpenAI: GPT-5.6 SolOriginal panel | 495/50099.0% · 97.68%–99.57% | 3,401 | 6,904 | $0.31276 |
| Qwen: Qwen3 30B A3B Instruct 2507Original panel | 495/50099.0% · 97.68%–99.57% | 1,830 | 3,262 | $0.00849 |
| Qwen: Qwen3 8BOriginal panel | 496/50099.2% · 97.96%–99.69% | 645 | 849 | $0.01668 |
| Qwen: Qwen3.7 FlashOriginal panel | 483/50096.6% · 94.62%–97.87% | 649 | 1,178 | —¹ |
| Qwen: Qwen3.8 FlashOriginal panel | 495/50099.0% · 97.68%–99.57% | 743 | 2,022 | —¹ |
API costs are billed charges; local cost is an allocation estimate. ¹ Missing billing records are marked with a dash. Wilson intervals describe sampling uncertainty, excluding domain shift and benchmark exposure.
Downloads: 30-model JSON, CSV; three-model extension JSON, CSV. Comparison charts.
Confidence gate
The gate accepts SmallDecide’s answer when its maximum candidate probability is at least the fixed threshold. Otherwise it uses the hosted model. On this test, 490 of 500 queries pass locally; ten need a fallback.
| Fallback model | Correct / 500 | Paired changes | Cost / 1,000 | Median ms |
|---|---|---|---|---|
| Claude Fable 5 | 497 → 491 | 1 gained / 7 lost | $1.9394 → $0.0617 | 3,209 → 162 |
| GPT-6 Astra | 498 → 492 | 1 gained / 7 lost | $1.5839 → $0.0548 | 1,349 → 162 |
| DeepSeek V4 Pro (0813) | 499 → 492 | 1 gained / 8 lost | $0.4478 → $0.0395 | 3,294 → 162 |
These are replays of recorded requests. All 500 queries were sent to every API during measurement. Accepted queries use their local duration; escalated queries add local and hosted durations. The table estimates a serial gate’s cost and latency, without measuring live queueing or deployed savings.
Threshold and errors
The threshold, 0.8502035140991211, was selected on 200 validation cases before the original test. It accepted 193 validation cases with one error. The test accepted 490 cases with 8 errors: 1.63% accepted-case error, above the 1% validation target. The threshold was not retuned.
The eight accepted errors never reach the fallback. Even a perfect fallback would cap this replay at 98.4% accuracy. Improving that ceiling requires changing the local model or acceptance policy, then testing again on new data.
“Gained” and “lost” count individual cases the gate corrects or breaks relative to the direct API. The results do not establish equivalent accuracy. That claim would require a predefined error margin and a paired analysis with enough independent cases.
Gate method · 30-model gate results · Extension records and replay code
Cost & latency
The local run used 166.105 warm compute seconds per 1,000 decisions. For an hourly allocation h and inference utilization u:
Fallback cost uses the actual charges for the ten escalated queries, scaled to 1,000 decisions. Multiplying average API cost by 2% would miss differences in their token bills. At 10% utilization, the allocated local cost is ten times the fully utilized estimate.
Excludes setup, operations, training, and error remediation on both paths; also excludes local model loading and API purchase fees and tax.
Timing conditions
Local timings include tokenization and synchronized warm inference. API timings include network travel, provider scheduling, and generation. Provider GPU types are undisclosed, so this is not an M4 Max versus B200 hardware comparison. Brief ERP probes also ran on the workstation during part of the three-model extension.
With 98% of queries accepted locally, the gate’s p95 can fall entirely within the local path and hide fallback delays. A production test should measure fallback latency, p99, queueing, and cold starts under a stated arrival schedule. Request latency alone does not establish throughput.
Candidate count
| Candidates | Median latency | Candidate-token evaluations |
|---|---|---|
| 2 | 18.9 ms | 244 |
| 8 | 21.1 ms | 976 |
| 32 | 51.2 ms | 3,948 |
| 128 | 192.7 ms | 15,908 |
| 255 | 374.3 ms | 31,910 |
Each candidate repeats the state in its prompt. Compute therefore grows with the candidate count even though no output tokens are generated. These H100 timings use different prompts from the M4 Max test and cannot give a hardware speedup ratio.
Model
The backbone is Qwen3-0.6B, with graph modules after layers 9 and 18. Each candidate is encoded with the state and question. Its score is the margin between the Yes and No logits, normalized using a temperature fitted on validation data.
Choice returns the highest-probability candidate. Score returns the expected zero-based rubric index, Σi i · pi. Noul returns the probability of the positive statement.
Confidence
The API’s confidence field measures distribution concentration: 1 − H(p) / log K. It is not a calibrated probability of correctness. The routing gate uses maximum candidate probability, a different quantity.
Graph modules
Each module has 16 learned nodes of width 128, four incoming neighbors per node, and two message-passing rounds. They read prompt representations and add a residual to the final decision token. The nodes are latent slots, not extracted business entities. Graph parameters and rank-16 LoRA adapters were trained together and merged into the release.
Removing graph message passing left greedy success rates unchanged in the reported native-environment tests. That ablation did not demonstrate a benefit from message passing.
Code: candidate scoring, API outputs, graph modules (repository access required). Ablation results.
Test setup
The release report covers 6,217 test and held-out decisions across 23 task families, including identified regression cases reused from earlier evaluations. This is a coverage count, not a pooled accuracy estimate.
Training and splits
Training used 28,234 prepared decisions from public datasets, verifier-labelled tasks, and redacted local-agent execution examples. Validation selected supervised step 1,400 and RL update 10; separate validation groups fitted temperatures. Split sizes and source hashes are in the release report. Private session rows are not redistributed.
The routing study used 200 validation and 500 test queries, excluding source training text, prepared local split groups, and CLINC groups from the release evaluation. Each query contains the correct intent and seven randomized distractors. SmallDecide was trained on CLINC training data; this test measures eight-way in-scope selection, not full 151-class classification or out-of-scope detection. Foundation-model pretraining exposure is unknown.
API conditions
- Original panel · 30 models
- Reasoning disabled where supported. One attempt per case, four concurrent requests per model and 24 globally. Default provider routing. The protocol records unavailable entries and the preflight amendment.
- Extension · three models
- Fable 5, GPT-6 Astra, DeepSeek V4 Pro. Low reasoning effort, 4,096-token reasoning/output cap, four concurrent requests globally. Reuses the already-inspected test set; exploratory results.
All APIs were asked for one compact label. Exact intent-name normalization was accepted across models; numeric-key compliance was recorded separately. Failed requests count as incorrect.
Original protocol · Splits and records · Extension protocol · Release report data. CLINC: Larson et al. (2019), CC-BY-3.0.
Factorio
Two identical circuit factories receive twelve scripted proposals at the same ticks. One executes all of them; the other executes only SmallDecide’s allow decisions. The game’s furnaces, assemblers, belts, and inserters produce the circuits. The runner inserts no finished circuits and does not write production values.
Watch the production shift 1 min 41 sec
SmallDecide blocks pole removal before a backup connection exists and allows it after installation. Final dispatch: 583 circuits with review, 482 without. The reviewed factory keeps its 200-plate maintenance reserve and first marks the shipment complete with 558 circuits loaded.
What the model receives
The adapter traverses power and belt connections and supplies textual preconditions. SmallDecide answers action-specific questions over those facts. It does not discover connectivity, plan the factory, or read video frames. The baseline is an execute-all script, not another model.
The film replays recorded decisions at 4× game time. Every intervention’s before/after state and tick matches the original run. The scenario was developed with the model; it is not a held-out reliability test. The first run’s mistakes and subsequent input probes are included in the downloads.
Original trace · Replay states · First run · Source · Reproduction
Audit tasks
In the release’s session execution forecast task, the default threshold detects 6 of 142 failures (4.2% recall). Always predicting success would score 89.0% accuracy. For this task, fault recall and false alerts are more useful than aggregate accuracy.
| Task | Initial correct / 48 | Best prompt / 12 cases |
|---|---|---|
| Completion evidence | 41/48 | 91.7% |
| Authorization scope | 28/48 | 75.0% |
| Expense policy | 28/48 | 66.7% |
| Workflow continuation | 36/48 | 75.0% |
Each task’s 48 cases are twelve wordings with four identifier variants, from six paired scenario families. A four-prompt search produced no task meeting the frozen 95% selection target. The separate holdout was left unevaluated.
Proposed follow-up
The next protocol would hold out scenario and template families, include deterministic rules and cheap-model gates, and freeze thresholds before test. It would measure fault recall, false alerts, unresolved cases, paired errors, total cost, and p95 workflow delay. Its targets of 99% fault recall and ≤5% false alerts have not been achieved.
Even those targets can create a large review queue. At an assumed 1% fault rate, 10,000 events would yield 99 caught faults and 495 false alerts: 16.7% alert precision. This is a hypothetical calculation, not a measured SmallDecide result.
Exact thresholds, booleans, and inventory rules also need a deterministic-code baseline. The Factorio adapter already computes connectivity; a useful follow-up must isolate what the learned reviewer adds.
Failure confusion matrix · ERP inputs and results · Proposed protocol
Reproduce
The extension bundle includes inputs, local and hosted records, manifests, the parser, and replay code. The command below verifies hashes, recomputes the validation threshold, and checks direct and gated metrics. It needs Python 3.10+ and httpx; no weights, credentials, or paid API calls.
curl -fLo routing-evidence.zip \
https://smalldecide.com/evidence/routing-frontier-20260917/evidence.zip
unzip -q routing-evidence.zip -d routing-evidence
cd routing-evidence
python3 -m venv .venv
.venv/bin/python -m pip install httpx
.venv/bin/python tools/reproduce_routing_frontier.py --source .
Rerunning inference requires the checkpoint and a compatible runtime; hosted calls also need credentials and a spending limit. Factorio requires a licensed 2.0.73 installation. Its archive contains no game assets or weights. Shorter examples are also available.
Checkpoint & sources
Routing and Factorio results use the v0.3.0 checkpoint with SHA-256:
ec83247177a6919c17944093fdd02d3258906dc809c16b326e0dfd879275efeb
The public-data checkpoint (f64a7763…) has a separate report. Its results are not included here. Tables are generated from source JSON; the build verifies source and media hashes.
Source hashesRelease dataRelease PDFPublic-run PDF
Evidence downloads are public. Repository and checkpoint downloads require GitHub access. Data attribution and code licensing are recorded in the bundles. Factorio is made by Wube Software; this is an independent demonstration.
Related: training and evaluation guide.