vastops: a self-hardening pipeline for training small models on rented GPUs
Status: living document.
Grounded: 2026-09-15, against the ledger and eval data in this repository (ledger events through 2026-09-15). Eight of the substrate table's ten rows are no longer typed: scripts/ledger-stats.py derives them and states every definition they rest on, and scripts/update-whitepaper-stats.py -- check fails if this document has drifted from the ledger. The two gate rows are the exception, and are counted by hand, because gate verdicts are not yet ledger events.
Citations: verified against their sources 2026-07-14, extended 2026-09-09 (see References).
Abstract
vastops is a small Python framework that turns the vast.ai GPU marketplace into a dependable training substrate for one operator. It wraps the vendor CLI with guardrails, verification, monitoring, and a cost ledger, and it wires the result into a deploy gate that refuses to ship a model that hasn't beaten its baseline on held-out data. Sixteen model versions have gone through it across fourteen active days, for a total marketplace spend of $41.62 with every boot failure and dead host included: nine cleared their gates, six were refused by their own gate, and one targeted a base its serving platform turned out not to host. The refusals are the point. One was diagnosed the same evening and its successor shipped through the identical bar before midnight; two more were caught by exact canaries that passing statistical aggregates had hidden -- once on an edge-serving candidate, once on a fine-tuned encoder whose validation numbers were nearly perfect. Four of the shipped models do production work today (asynchronous internal tooling and routed product features, not public low-latency serving), and two of those supervise the training pipeline itself.
Two things have changed shape since this document was first grounded. A second marketplace shipped, which turned the vendor seam from a design intention into something that has rented, trained, and been billed under the same guardrails. And the deploy gate stopped being a convention a project echoed into a log and became a contract the framework reads and exits on, where a metric an eval failed to produce counts as a failure rather than as silence. Both came out of the same rule: whatever the pipeline relies on, it should be able to check.
The interesting property is not any single feature. It is the loop: every operational failure becomes code the same day, every research finding becomes a recipe change or is discarded, and the models the pipeline produces become the monitoring layer for the next training run. The pipeline gets harder to break every time it breaks. It exists to serve a bet: that production AI decomposes into narrow, high-volume tasks where a fleet of small specialists beats one giant generalist on accuracy, latency, and cost at once, and that the winner of that game is whoever owns the cheapest honest model factory.
The problem
Marketplace GPUs are cheap and hostile. A single RTX 3090 rents for eleven cents an hour, which prices a 4B-parameter QLoRA fine-tune at roughly a quarter. But the hosts are consumer machines run by strangers: they misreport hardware, freeze mid-docker-pull, ship drivers too old for the image you asked for, and keep billing after your training process has died. Rent one naively and you learn that the advertised price was the easy part.
Meanwhile the case for small fine-tuned models keeps getting stronger. For narrow, structured tasks, a fine-tuned 0.6B–8B model routinely beats far larger prompted models, GPT-4-class systems included in the published record, at a fraction of the latency, on hardware you own. Our own measurements below bear this out on every task we've tried. The bottleneck is not model quality and not GPU cost. It is everything around the training run: provisioning, verification, monitoring, artifact integrity, evaluation honesty, and deployment safety.
vastops exists to make everything around the run mechanical.
Architecture: policy around a replaceable vendor core
A stdlib-only Python package wraps the official vastai CLI (JSON mode, retry on transient failures) rather than the raw HTTP API, so vendor API drift is the vendor's problem. On top of that sit guardrailed search and provisioning, rsync-based code sync, tmux-managed training sessions, regex-parsed progress reporting, an idle watchdog, and a lifecycle ledger that records every create, stop, and destroy with price and reason. Per-project behavior lives in a vastops.toml checked into each training project. Guardrails live only in the operator's config file: a project can define its image and sync paths, but it cannot raise a price ceiling or lower a reliability floor.
The load-bearing decision in that stack is where the vendor stops and the policy starts. Exactly one module talks to the marketplace, and it does so through the vendor's own maintained tool; everything above that seam reasons in provider-neutral terms about offers, instances, prices, and lifecycle states. The guardrails, the ledger, the delivery verification, the watchdog, the self-healing loop, and the deploy gate never mention vast.ai. Neither does a project's vastops.toml, which describes a training run (image, disk, sync set, progress regexes), not a cloud. The practical consequence is that the provider is a dependency while the policy layer is the product: swapping or adding a marketplace changes one adapter, and every hard-won rule about money, verification, and honesty carries over intact. The ledger and the lessons file outlive any vendor relationship, which is exactly the property an operator wants from a system that remembers what things cost and what went wrong. That seam has since been tested rather than asserted: a RunPod adapter shipped and has rented, trained, and been billed for. The claim it discharges is narrow but real, because the adapter had to straddle two of RunPod's three API surfaces (REST v1 for pod lifecycle, the only one exposing a public IP and therefore rsync; REST v2 for the GPU catalogue, the only pricing query) without any of that leaking above the seam. The policy layer above it did not change. What the adapter also proved is that a provider brings its own lies: RunPod reports a stopped pod's full compute rate as though it were still burning, returns business-rule rejections dressed as HTTP 500, and turns stop-then-start into a trap where the pod can come back with zero GPUs attached while its volume bills at double. Each of those is now an adapter-level correction rather than folklore. One adapter is not portability; it is evidence that the seam is in the right place.
Trust nothing, verify everything
The design assumption is that marketplace hosts are untrusted until proven otherwise, every boot:
- GPU delivery is verified, not assumed. After an instance reaches
running, the pipeline asserts the advertised GPU count via nvidia-smi and then runs a CUDA compute preflight:torch.cuda.init()plus a bf16 matmul. A host that passesnvidia-smibut cannot actually compute (stale driver against a new CUDA image, the failure that once idle-burned 105 minutes) is destroyed and replaced in about ten seconds.
- Boot is fail-fast.
BootProgresswatches the host's status messages during image pull; a frozen status line for six minutes, or a known-fatal pattern, kills the attempt. Before this existed a bad host cost 25–30 minutes of poll-waiting; now it costs three or four. Neither SkyPilot nor dstack shipped an equivalent as of July 2026, because both assume cloud-grade hosts -- and the honest comparison runs the other way too: those tools are better at what they actually do (multi-cloud placement, spot orchestration, managed jobs across friendly providers). What they assume away is exactly what this marketplace sells. vastops is the complement, not the replacement.
- Bad hosts are remembered, and so are good ones. A boot or verification failure auto-quarantines the machine id for a cooldown window, so the retry loop cannot re-pick the same dud; persistent offenders go to a permanent blacklist. On the positive side, every rental outcome feeds a per-machine reputation: a time-decayed Beta posterior over boot-and-deliver success, graded into coarse letter buckets and blended into offer ranking as a soft multiplier. Quarantine stays the hard filter and the grade only reorders, never excludes; failures our own code caused barely touch a host's score, and a changed hardware fingerprint behind a stable machine id resets its reputation. The marketplace's own reliability number measures host uptime, which is why a host can throttle its GPU 22% and keep a clean score; this ledger is renter-side and measures delivery, which makes it much harder to game than an uptime score -- though not impossible: a host can pass a ten-second preflight and degrade later, which is why utilization sampling continues for the life of the rental and why the grade reorders but never overrides the hard floors.
- Offer selection is value-based. Search sorts by DLPerf per dollar under floors for reliability (0.98), download bandwidth (400 Mbps, learned after three hosts sat cold-pulling a 5GB image on slow links), and CUDA driver recency.
- Artifacts are verified end to end. The exported GGUF carries a provenance manifest: sha256 of the artifact, the dataset, the config, and the git revision that produced it. The pull is resumable and retried (one 2.4GB artifact took seven attempts off a degrading host); destroy happens only after the local hash matches the manifest. A stale-adapter guard refuses to export an adapter older than the training-start marker, which kills the "shipped yesterday's weights" class of mistake. That discipline started as a per-project shell script and is now a framework command:
vastops pull -- verifyhashes the declared artifacts on the instance, hashes what landed, and reports missing and corrupt files separately from merely-extra ones. rsync exiting zero says the transport did not error. It does not say the bytes match, and on a host that can die mid-write those are different claims.
Why not SkyPilot, and what it does better
The question every infrastructure reader asks first deserves its own answer, re-checked September 2026 against those tools' current releases. Two things changed since the first version of this table and both cut against vastops, so they lead. SkyPilot ships a vast.ai backend (March 2026) and a RunPod one, so "we drive a marketplace nobody else drives" is simply no longer true. And both SkyPilot and dstack now have hourly cost ceilings (max_hourly_cost, max_price), so a row that used to read "hard price guardrails: no / no / yes" was overclaiming. The surviving distinction is narrower and worth stating precisely: their ceilings live in the same YAML a project author edits, ours live in the operator's config and a checked-out project is forbidden from touching them. That is a real difference for a framework whose threat model includes its own project files, and a small one otherwise.
| Capability | SkyPilot | dstack | bare vastai CLI | vastops |
|---|---|---|---|---|
| -- - | -- - | -- - | -- - | -- - |
| Multi-cloud placement and spot orchestration | strong | strong | -- | none (by design) |
| Managed job queues, many-user teams | strong | strong | -- | none (by design) |
| Boot-progress inspection with frozen-pull fail-fast | assumes cloud-grade hosts | assumes cloud-grade hosts | -- | yes |
| Post-boot CUDA compute verification (not just nvidia-smi) | no | no | -- | yes |
| Renter-side host memory (quarantine, delivery record) | no | no | -- | yes |
| Cost ceiling | yes (max_hourly_cost) | yes (max_price) | -- | yes |
| Ceiling a project's own config CANNOT raise | no | no | -- | yes |
| Append-only cost ledger surviving instance destruction | no | no | -- | yes |
| Artifact-hash-before-destroy invariant | no | no | -- | yes |
| Model deploy gate (candidate tags, canaries, refusals) | out of scope | out of scope | -- | yes |
What survived re-checking, and it is the least contestable part: nobody else verifies the host from the renter's side. SkyPilot's vast.ai integration adds a datacenter_only filter, which excludes consumer hosts rather than testing them; the community's answer to "did I get the GPU I paid for" is still to SSH in and run gpu-burn by hand. No orchestrator surveyed runs a post-boot compute micro-benchmark, a PCIe-link check, or keeps a renter-side delivery record. Artifact-hash-before-destroy likewise turned up nowhere in the competitive set.
One number in this paper has no external corroboration and should be read that way: the 64.8% never-delivered rate is first-party, from our own ledger, and a search for a comparable published figure for marketplace GPU rentals found none. The academic 2026 GPU-reliability work studies operator-owned production clusters, which is a different population. The figure is unrebutted, not confirmed.
The honest reading of that table is that these are different tools. SkyPilot and dstack are schedulers for people with many jobs across trustworthy clouds, and they are better at that than vastops will ever be. vastops is a survival-and-honesty layer for one operator renting from strangers, plus a promotion discipline no scheduler attempts because it is not a scheduling problem. The plausible objection "SkyPilot with a preflight hook" founders on control-loop ownership: an aggressive fail-fast watchdog stacked on another orchestrator's retry logic gives two recovery systems fighting over one instance, which is why the roadmap's next backend is a thin direct adapter under conformance tests rather than a layer on a layer.
Money is a first-class failure mode
The ledger is append-only and survives instance destruction, so "what did we run and what did it cost" always has an answer. The watchdog samples GPU utilization on a cron and stops anything idle past a grace period; a stopped instance bills cents of storage instead of dollars of GPU. Stop and destroy are deliberately asymmetric: automation may stop, only the operator destroys, because a wrong stop costs cents and a wrong destroy loses a training run. Duplicate-instance leaks are closed structurally: a create that times out is reconciled by label before any retry, so a lost API response cannot leave an unrecorded instance billing forever. And the pipeline script itself cannot leak a rental by dying: an exit trap stops the instance on any abort, an operator's Ctrl-C included, and disarms only after the intentional destroy.
The efficiency canon is ledger-grounded rather than aspirational. Reading the actual event history showed that provisioning, not training, is the bottleneck (one campaign burned 22 create/destroy cycles against 33 minutes of actual training), which produced concrete rules: size the GPU pool to the model, probe the market before renting, batch multiple quick fine-tunes onto one verified rental, and run the prompted baseline on the local inference box in parallel with provisioning.
The deploy gate, or: how a model earns its tag
Every model ships through the same ceremony, and the ceremony has teeth.
1. A prompted baseline must lose first. Before training, the task's eval harness runs against a much larger local model with a well-crafted prompt. If the fine-tune can't beat a free prompt, it has no reason to exist. This is measured, not asserted.
2. The bar is judged on the worst seed. Evaluation runs on two held-out seeds and the gate reads the worse one. A lucky seed cannot ship a model.
3. Candidates never touch the live tag. The new artifact deploys as 4. The bar never relaxes to go green. Loosening a threshold is a policy change requiring the operator, not a fix. 5. The gate is a contract the framework can read. It used to be a convention: the project echoed This is not theater, and it has now refused two models. ops-doctor v3 improved on the metric its enrichment targeted while silently losing 45 points on ssh-unreachable hard negatives; it evaluated as a candidate, failed head-to-head against v2, and never reached the live tag. The postmortem (checkpoint selection by eval-loss, which does not track label accuracy) became a recipe rule that v4 shipped with. sign-fielder v1 beat its prompted baseline by 2.5x and the gate refused it anyway, 2.6 points short on signature recall. Triage traced both failing metrics to a single cause: the output format nested the field grammar's braces inside JSON strings, and the six garbled replies held exactly the signatures the bar missed. v2 switched to a collision-free line format and cleared the identical bar the same evening. A gate that only ever passes things is decoration; this one earns its keep by saying no. A skeptic will notice that the early refusals were both fixed within a day and read the gate as a delay, not a barrier -- but the gate's product is the forced diagnosis. Both root causes (checkpoint selection by eval-loss, a brace collision in the output format) would have shipped silently without it, and one of the later refusals has not been fixed at all: the 3B edge adapter stays refused, because the honest answer so far is that the capacity isn't there. The ceremony also grew this week. Ten canary inputs with exact expected outputs run before the statistical eval, because a rare catastrophic regression does not move an n=86 aggregate but trips an exact expectation immediately. That stage has since earned its keep for real: an edge-serving candidate passed the statistical bar outright (F1 0.924, signature recall 0.979) while failing the same multi-party canary four times out of four. The aggregate had averaged the miss away; the canary refused the model. The canaries are also the only stage computed entirely on the operator's own box against exact expectations, which makes them the de facto anti-tamper check on artifacts pulled from untrusted hosts -- a job ten inputs is thin for, and one that costs nearly nothing to grow for grammar-constrained tasks. A hand-labeled out-of-distribution slice reports alongside the bars (advisory until calibrated, then gating). And promotions over an existing incumbent now have a paired-statistics stage available (McNemar on discordant items, bootstrap intervals on per-document deltas) that returns three verdicts, better, worse, or unresolved, because at these sample sizes a naive comparison routinely reports noise as a winner [13]. Unresolved means the absolute bars decide; the paired test only has authority when it has resolution. "Dependable" is a measurable claim, so the substrate numbers come first. From the ledger, 2026-07-11 through 2026-09-15: Generated by Roughly two in three creates died before any training happened, and on the two documented platform-wide bad days the failure rate hit three in four. That is not the pipeline failing; that is the marketplace being what it is, made cheap: the number to judge is the cost of those deaths, $4.96 across the whole history at under thirteen minutes per bad host, because detection is fail-fast rather than poll-and-hope. Those bad days are why a second provider went from queued to shipped, and RunPod rentals now sit in the same ledger under the same rules. All rows are held-out and gated on the worst of two seeds, except ops-doctor v4, whose promotion was a single-seed head-to-head against its shipping incumbent (n=300). "Baseline" unless stated is qwen3.5:35b-a3b given the same task with a carefully written prompt: a 35B-total-parameter mixture-of-experts that activates roughly 3B parameters per token, and the strongest generalist the 24GB serving box runs -- which is the honest constraint, because the practical alternative to each fine-tune is a prompt against this box, not a paid API. The baseline got the same output-format iteration the fine-tunes did (sign-fielder's prompted baseline was measured in both the JSON and line formats, and the better number is the one reported). To be precise about sizes: against the 0.6B fine-tune the baseline is 58x the total parameters; against the 8B it is 4x the total parameters but fewer active parameters per token. The honest reading is not "David beat Goliath" at every row. It is that a model trained on the task beats a much broader model prompted into it, whatever the parameter ratio, and the frontier-scale comparisons (GPT-4-class) come from the published record cited in the thesis section, not from this table. ¹ ops-doctor's reference is its shipping incumbent, not the prompted 35B -- by v4, the incumbent was the harder bar to clear. ² A perfect score on a simulator-generated eval means the eval has nothing left to teach; the real-corpus retrain replaces this number. ³ Measured with an embedding-similarity sweep. Validation is deployment-shaped for a template-driven product; the OOD slice is the number that keeps it honest, and the regex incumbent still beats the fine-tune on that slice. Three patterns repeat. The fine-tune beats the prompted baseline on the three axes the bars declare: task accuracy, the task's costly error, and latency. The training cost rounds to pocket change. And the asymmetric error that matters for the task (cutting a user off mid-sentence, calling a dead run healthy, missing a signature line) is a named, gated metric, not an afterthought. A fair objection to the table above: beating a prompted 35B does not establish that a fine-tune beats a good deterministic parser or a right-sized encoder. So the newest task ran as a four-way fair fight on identical rows: PII span detection for the e-sign product, on a dataset whose training and held-out format pools are deliberately disjoint. Nobody wins all three columns, and that is the finding. Rules are perfect on the surface forms their author has seen and structurally blind past them. The fifteen-cent encoder dominates the deployed distribution -- and scores zero on the highest-severity classes in unseen formats, a perfect surface-form memorizer exposed by the format-disjoint split built to catch exactly that. Only the prompted generalist degrades gracefully under drift, and it is mediocre everywhere at twelve times the encoder's latency, with 5-9% invalid outputs. The gate refused the encoder despite its near-perfect validation numbers, on canaries; that refusal, of our own best candidate, is the strongest evidence in this paper that the ceremony measures models rather than flattering them. The production answer is a cascade, not a winner: rules for the known, a specialist for the deployed distribution once format augmentation closes its drift gap, and the generalist as the fallback for the genuinely new. The honest numbers ship too, and they drive retrains. ambie-eou v1 reported its external out-of-distribution slice at 80.9% (an independent team's test set, versus 99.3% in-distribution) and diagnosed why: inputs with no silence signal inherited a prior from an injected placeholder value. v2 trained with the silence field randomly absent, which is standard feature dropout, and re-measured both versions on the identical, honestly-reshaped slice: v1 scored 82.4% with 29.8% false interruptions, v2 scored 90.0% with 16.4%. The out-of-distribution number improved because it was reported instead of hidden. log-doctor's README carries a same-simulator caveat. sign-fielder's record carries two of its own: an advisory out-of-distribution slice (67.4% structural pass on independent upload-style documents, against a regex incumbent that scores 100% on what is effectively its own tuning set), and a measured 51% semantic-sibling rate between its train and validation sets -- quantified with an embedding-similarity sweep and interpreted in the open: validation measures in-distribution performance, which is deployment-shaped for a template-driven product, and the out-of-distribution slice is the number that keeps it honest. Eval theater is a failure mode this pipeline treats like any other: something to detect and close. The comparator question above asks which model wins. A different question sits underneath it: which route to a specialist wins. A note circulated in September 2026 argued the fashionable one, structured pruning of a 7-8B model followed by knowledge distillation, and it is worth being precise about why this pipeline does not take that route. Five routes exist. Fine-tune the smallest architecture that already matches the output shape. Fine-tune an already-small current general model. Put a LoRA adapter on a base that is already resident. Quantise. Prune a bigger model and distil it back. Six independent expert reviews and five literature passes, run against this question in September 2026, put prune-and-distil last on all six, and the published evidence is what makes that more than opinion. Token-matched, pruning loses: at equal total token budget, from-scratch training beat coarse structured pruning of Llama-3.1-8B by 0.40 perplexity on WikiText-2, and the paper's own reading is that architecture transfers while weights largely do not [14]. On the canonical demo task, aggressively pruning GPT-OSS-20B's experts to about 5B and recovering lands competitive with rather than ahead of NLLB-200-3.3B and TranslateGemma-4B [15] -- while Three specific claims in that note are wrong in ways worth recording, because they are the kind that survive by sounding plausible. Attention probabilities are softmax-normalised and sum to one for every query, so a head's average weight is about one over the context length even when the head is vital; "heads with near-zero attention weights are dead" is not a measurement. Under grouped-query attention, dropping a query head does not remove its shared key-value projection or shrink the KV cache, and NVIDIA's own Minitron ablation kept attention intact and pruned only MLP and embedding channels because that scored better. And the memory wall in logit distillation is not two sets of weights, it is the None of that is a reason to take the conclusion on trust, so The split that will decide it is It ran on 2026-09-09. Every arm trained on the same corpus at the same budget (1,200 optimizer steps at an effective batch of 32, so roughly 38,400 examples each), scored by the same evaluator on the same splits. The whole matrix, ten scored arms, cost $2.63 of rented GPU: three creates on RTX 5090s, of which two never delivered a working GPU and one of those had to be destroyed by hand after an external Seven of the ten scored arms are shown; the remaining three ( The thesis held. A fine-tuned 0.6B beat a prompted 4B, a model 6.7x its size, on every split, and the paired McNemar on identical items says BETTER on val (+36/-21) and BETTER out of distribution (+44/-24), not noise. That is the controlled version of this paper's central bet: same family, same corpus, same evaluator, same output contract, with specialisation the only variable. Five things the bench said that we did not expect, and three of them corrected a claim already made in this paper. Zero-shot ranking does not predict fine-tuned ranking. Vocabulary trimming is not the cheap lever it is described as. The idea is sound on paper: 22,992 of 151,643 tokens are ever used by this corpus, and the embedding is a quarter of a 0.6B. Measured, trimming without a recovery pass took the 1.7B from 0.874 to 0.176, with output validity collapsing to 0.065 and latency rising tenfold because the damaged model never emits a stop token. The cause is the failure mode the design predicted and the write-up still underweighted: the kept set comes from the training corpus, so eval prompts contain tokens that become UNK on input, not merely unemittable on output. Trimming plus a recovery pass may well work; trimming alone is destructive. Quantisation trades speed for memory, not speed for size. Prefill does not dominate; decode does, by thirty times. This bench was built expecting a few-thousand-token JSON Schema prompt to dominate a few-dozen-token answer. Prefill measures 27-54ms against 1131-1680ms end to end. Shorter prompts will not help this task; fewer output tokens or faster decoding will. Fine-tuning bought format compliance for free. And the pruning arm never ran, because the planner refused it. Asked for 350M from a 1.7B parent, it reported that no honest plan reaches that target without cutting the MLP below a quarter of its width, and stopped. The guard was written that morning against a different model and fired correctly on a request made in good faith hours later. A second experiment, 2026-09-10, on the same harness. It exists because the first one measured the wrong resource. VRAM in the table above is what a model costs to hold. What bounds a serving fleet is what a request costs to run, and that is KV cache. Qwen3-0.6B and Qwen3-1.7B both spend 112 KB per token (28 layers, 8 KV heads, head_dim 128), so on a 32GB card the 0.6B's 2.2GB weight advantage is worth about 8% more concurrent requests rather than the 2.9x its parameter count advertises. Against the 4B the honest multiple is roughly 1.7x, not 5.2x. Anyone sizing a fleet from the weights column would overestimate it badly. KV cache scales with prompt length, and for this task the prompt is mostly JSON Schema describing tools the model has seen tens of thousands of times in training. So: is the schema in the prompt buying anything the weights do not already hold? Five arms, $1.25, same seed and budget as the control. It is buying something. Every arm is WORSE on a paired McNemar over identical items, none of the gaps is near the bench's ~5 point resolution, and the gate is set at. It would not ship. Three things fall out that the token counts alone do not predict. The catalogue is not in the weights. Truncating the prompt on an already-trained model costs 8.8 points at retraining at the short prompt recovers only part of that. The model reads its schema at inference; fine-tuning on a fixed tool set does not memorise it. Damage is wildly disproportionate to what you save. Dropping the prose descriptions costs 209 tokens for 7.1 points. Dropping the argument names and types costs 34 further tokens for another 11.9. Per token removed, the second cut is about 10x worse. The prose is cheap; the signature is load-bearing. A plan built on "the prompt is mostly boilerplate, cut it" would have cut the wrong half. And it buys no latency at all. 521 / 517 / 536 ms p50 across a 2.9x range of prompt length, prefill flat at 14-15ms. That is the overhead floor above, confirmed from the opposite direction: shrinking the model did not move the clock, and neither does shrinking the prompt. So the real trade is 3.4x concurrency for 7.1 points of accuracy, and at that exchange rate renting 3.4x the GPUs is cheaper than losing seven points. The lever that does survive is orthogonal and unglamorous: an fp8 KV cache halves KV per token for one flag, no accuracy cost, and no experiment needed to know it works. The failure mode this experiment was really testing for is the one where an elaborate idea is preferred to a boring flag, and the boring flag won again. The latency column measures the harness at least as much as the model. Every p50 above comes from HuggingFace that wall clock (42 output tokens against 31ms of prefill). Check it against the hardware: 1.8 TB/s, so reading every weight once takes 0.67ms -- against a measured 30ms per token, which is 45x off the roofline. The 1.72B model decodes at 29.1ms/token and the 0.60B at 30.0ms/token, near enough identical despite a 2.9x difference in parameters. That is the signature of a fixed per-token overhead -- Python loop, sampling, cache management -- setting the clock while the GPU idles. So these numbers rank the arms fairly against each other, because every arm paid the same overhead, but they must not be read as "a smaller model is faster": at this size class, on this serving path, it mostly is not. A real latency claim needs a serving runtime that removes the overhead floor (vLLM, llama.cpp, TensorRT-LLM), and we have not run one. And the overhead floor is not even a constant. on 2026-09-10, same weights, same dtype, same GPU model, output length within 10%: it came back at 13.5ms per token against 29.1 the day before, a 2.2x swing attributable to a different physical host and to 5.17. So the latency figures are comparable BETWEEN arms of one run, which is what the table uses them for, and are not portable across runs or quotable as absolutes. Note the contrast with the accuracy numbers, which reproduced on that same re-run to within the bench's own resolution (val 0.8875 against 0.8736). The eval is stable; the clock is not. And VRAM is not concurrency, which the prompt-length experiment above measures rather than asserts: the weights column is what a model costs to hold, not what a request costs to serve. The baseline is zero-shot, not few-shot. A few-shot prompted 4B might close some of the two-point gap, and this experiment does not test that. It is the same gap the 2026 literature has, and naming it is better than letting the headline travel further than it earned. It is also one task, on a public corpus old enough to be in every pretraining mix here. The small val-to-OOD gaps argue against memorisation dominating, but do not eliminate it. And the arms differ in learning rate as well as in method, so a result inside a couple of points carries a hyperparameter choice with it, which is what the paired test is there to stop anyone over-reading. This is the honest shape of a claim about method: not "we chose the best route", but "here is the bench on which the routes are ranked, and here is what it says". When a model leaves the training box, the rules go with it. The production field-detection model now serves over two paths. The primary is an authenticated tunnel from the product's edge functions to the home inference box: the exact gated artifact, same weights, same decoding configuration, behind mutual service-token auth. The fallback tier is a LoRA adapter hosted natively on the edge provider's own base model, retrained from the same dataset because an adapter is dimension-specific to its base. The pipeline refuses to treat these as the same model: the tunnel serves what the gate approved, while the edge sibling must clear the identical bars on its own edge-served outputs before the router may fall back to it. That rule has done real work. The edge sibling has been refused three times, each for a different reason. First the canary caught the multi-party miss a passing aggregate had hidden (the incident in the gate section above). Then the provider's documented LoRA rank limit of 32 turned out to be a real limit of 16, discovered when a rank-32 adapter uploaded cleanly and then failed every inference call; a local validator now rejects over-cap adapters at upload so no future GPU run re-buys that lesson. Third, with the strongest recipe the platform legally serves (rank 16, all linear modules) and doubled multi-party training data, the 3B base still missed the same signature block every time: a genuine capacity wall, not a data problem. Trust-nothing turns out to apply to serving vendors as much as to marketplace hosts, and the same loop metabolizes both: undeletable fine-tunes, required config fields the trainer omits, and a roughly 40% transient failure rate on freshly-uploaded adapters are all encoded as code, not folklore. The capacity lever worked. The same recipe retargeted at the provider's strongest tunable base, a 7B at 2.3x the parameters and about seventy cents of GPU, went ten-for-ten on the canaries including both historical failers, cleared the statistical bars on the worst seed, and in the paired head-to-head against the exact-model tunnel incumbent came out measurably better on document F1 (bootstrap CI95 [+0.018, +0.056], entirely positive), with the signature discordants unresolved at n=2, which is what honest resolution reporting looks like at that count. The edge tier has its fallback; go-live is one flag-gated router line and a secret, on the operator's schedule. Three model sizes, five dead attempts, and one vendor documentation error separated the idea from the artifact, and the gate is what kept every intermediate from shipping. The results table is not a run of luck. It is the local evidence for a thesis this operation has held since the AMBIE autonomous-ASR design: production AI work decomposes into narrow, stable, high-volume tasks, and for those tasks a small model trained on exactly that task beats a giant general model on accuracy, latency, and cost at the same time. The AMBIE spec drew it before this pipeline existed: a router in front of a fleet of specialized recognizers, one tuned for industrial noise, one for medical vocabulary, one for legal, with the general model as the fallback rather than the workhorse. vastops is that idea built out as infrastructure, one specialist at a time. The external evidence has caught up, and it needed refreshing: the citations this paper leaned on were measured against GPT-4, a baseline two generations stale. The current one is better for us. "Sub-Billion, Super-Frontier" (June 2026) fine-tuned a Qwen2.5-0.5B on relation extraction and beat zero-shot GPT-5.4 (0.83 vs 0.69 micro-F1) and Claude Sonnet 4.6 (0.66), rising to 0.92 against 0.83 in the literary domain. An in-domain RoBERTa baseline also beats both frontier models, so the authors attribute the win to task adaptation rather than to anything about model type [14]. On tool calling specifically, fine-tuned 3B variants match 7B ones, and fine-tuning lifted tool-selection recall 45 points [15]. NVIDIA's position paper still measures serving a 7B model at 10–30x cheaper than a 70–175B one and estimates 40–70% of three real agent frameworks' calls are narrow enough for a specialist [1]. Predibase's LoRA Land remains the largest breadth result, at 310 adapters over 31 tasks, but its GPT-4 comparison should now be read as historical [2]. A fine-tuned DeBERTa-large, roughly a third of a billion parameters, scores 87.8% on legal clause classification where GPT-4 zero-shot scores 67.2% [3]. Meta's Llama Guard 3, an 8B guardrail model, scores 0.939 F1 on its hazard benchmark where GPT-4 scores 0.805 [4]. LiveKit runs a half-billion-parameter turn detector (a Qwen2.5-0.5B fine-tune) in production on CPU, and a single version bump cut its false interruptions by ~39% [5]. Our own table sits comfortably in this company -- as consistent-with, not proof-of: every fine-tune beat the strongest prompted model our box serves on its declared axes, including a 0.6B that answers in 140ms where a model with 58x its parameters took half a second to be wrong more often. The part of the thesis that sounds like arithmetic but changes the architecture is VRAM co-residency. A 4-bit 4B model occupies about three gigabytes; a 24GB consumer card holds a handful of distinct specialists at once, and adapter-based serving pushes the marginal specialist toward zero, since one resident base model can host thousands of task adapters that page in per request. That was a research claim when this paper first cited S-LoRA [6]; by 2026 it is shipped infrastructure -- vLLM's multi-LoRA is generally available and runs behind AWS SageMaker and Bedrock, benchmarked with eight adapters resident at once [18]. The operational friction that remains is cold-loading a new adapter and mismatched ranks across adapters, not the idea. One box on a home LAN already serves this operation's whole fleet: two ops models, a turn detector, a triage model, and a field detector, with room to keep going. The giant-model equivalent of that fleet does not fit on the card at any quantization, and renting it means paying know-it-all prices for skills the task never uses. This is worth distinguishing from mixture-of-experts, which routes among experts inside a single jointly-trained network and ships as one deployment. A fleet is looser and more operable: each specialist is trained, evaluated, gated, promoted, and rolled back independently. When the field-detection grammar changes, the turn detector doesn't retrain, and a regression in one specialist cannot leak into another. The router is dumb and explicit instead of learned, which is precisely why it is debuggable. The loudest 2026 counter-signal is that OpenAI deprecated self-serve fine-tuning in May 2026, closing it to new organisations immediately and ending all self-serve jobs by January 2027. Their stated reason goes straight at this thesis: newer base models follow instructions and formats much better, prompt-based approaches are cheaper and faster, and they see fewer use cases that require fine-tuning [16]. That evidence is real and it is about a different thing. It concerns fine-tuning a closed-weight frontier model through somebody's API, where you pay a premium for a checkpoint you do not hold and cannot serve yourself. This pipeline fine-tunes open weights on rented hardware and serves them on a box in the house. The economics and the failure modes are not the same, and a reader should not let one stand for the other. But the underlying claim, that better base models erode the need for adaptation, applies to both, and it is the strongest argument against this whole enterprise. Three more erosion points, named rather than dodged: That last point deserves its inversion, because the evidence cuts the other way too. "The Constraint Tax" measured hard-schema constrained decoding on small models: validity rises from 61.5% to 100% while answer accuracy falls from 19.7% to 11.0%, and a Qwen2.5-1.5B on a tool-call task drops from 91.5% executable accuracy under plain prompting to 48.0% under a hard schema [17]. Constraining the grammar does not fix wrong answers; it makes wrong answers well-formed. And at the frontier, format validity and semantic correctness have plainly decoupled: a model can hit 99.97% JSON pass rate and 48.6% fully-correct responses. Our own numbers corroborate the useful half of that. The fine-tuned 1.7B tool-caller reached 1.0000 output-validity with no constrained decoding at all -- fine-tuning bought format compliance for free, without paying the constraint tax. Anyone reaching for a grammar to fix a small model's output should measure what it costs them in accuracy first. Finally, the honest gap: nobody has published a rigorous "we built the specialist and prompting won" postmortem with numbers. The 2026 case against is structural and economic, not an empirical failure. This paper will not invent one to knock down. The honest boundary: specialists win when the task is narrow, stable, and called often. Open-ended reasoning, low-volume work, and fast-shifting requirements still belong to generalists, where long context and in-context learning inject task knowledge without a retrain. The production answer is heterogeneous, and this pipeline's own monitoring ladder is the pattern in miniature: local specialists handle the routine calls for electricity, and a frontier model sees only the ambiguous residue. The real moat is not any single small model. It is the factory: a fleet strategy is only viable if training a new specialist costs an evening and a quarter, retrains are cheap enough to run on every material research delta, and a gate keeps each one honest. That is what this pipeline is for. This is the core claim, and it rests on four loops rather than one. Loop 1: every incident becomes code, the same day. The rule is that an operational lesson may not live in memory or in a doc alone; it must land as a commit that makes the failure structurally impossible or automatically handled. A sample from one week: a WAN flap false-killed a healthy run, so the watch loop got a six-cycle grace window and now distinguishes SSH-down from session-ended, corroborating against the marketplace API before declaring death. A shell one-liner bound its log redirect to the last command only, so a pip failure was invisible and healthy hosts got blamed; the train command is now group-wrapped, and that's a code-reviewed invariant. A broken direct port on one host's NAT produced an unreachable instance, so SSH and rsync both learned a proxy fallback, proven live the same week. The loop found a class it had not seen before on 2026-09-09, and the class is worth naming: a guardrail that fails closed, silently. Three separate filters were doing nothing, and each looked like a fact about the world rather than a bug. None of the three would have shown up in a test suite at 100% branch coverage, because all three were One day shows the loop closing in real time, twice. A variable-length batch OOM'd a 24GB card at step 55. The pipeline stopped billing within one poll cycle, kept the disk, and the run resumed on the same host twenty-two minutes later with a halved micro-batch. Then the incident itself got generalized into Loop 2: research runs on a protocol, not on enthusiasm. The first training run of each day starts with a fresh-knowledge pass: two or three cheap research agents check for new small-model releases, toolchain breakage, and task-SOTA movement, plus a thirty-second live market probe. The pass has explicit act thresholds (a candidate must plausibly beat the incumbent by two points or halve inference cost, otherwise it is logged and ignored), so research that doesn't change a decision is recognized as waste. When a delta is material, the affected shipped models queue a retrain through the same gate as everything else. This is how the pipeline absorbed an axolotl packaging bug, a torch pin conflict that silently killed three runs, and the switch to a 5GB platform-cached image after third-party registry throttling wedged four consecutive hosts in one afternoon. Loop 3: the pipeline's own products supervise it. The always-on monitoring ladder is deterministic detection first, model judgment second: thresholds and crash-signature regexes flag, a local triage model discards the benign, ops-doctor proposes an action, log-doctor reads the raw log tail when the regexes come up empty, and only the genuinely ambiguous cases escalate to a frontier model with a capped budget and a deny-by-default tool gate. ops-doctor and log-doctor were themselves trained by this pipeline, on data generated from its own failure taxonomy. Each new model that ships makes the babysitter better at catching problems in the next training run; the silent idle-burn a human once noticed after 105 minutes is now caught in about ten -- by the deterministic watcher, with a thirty-five-cent model reading the log tail to say why. An autonomous stop requires a deterministic detector to corroborate the model's opinion, so a small model's occasional overreach cannot kill a healthy run. The known blind spot here is circularity, and it deserves naming rather than narration around it: models trained on the pipeline's own failure taxonomy are exactly as blind as the taxonomy, and so are their evals. The corroboration rule contains the damage a wrong model opinion can do; it does not manufacture novelty detection. A failure class nobody has seen yet still depends on the deterministic floors, the frontier escalation path, and the operator -- which is why the escalation ladder ends at a frontier model instead of at the specialists. Loop 4: the tests police the tests. The package holds 100% branch coverage (704 tests at time of writing) as a hard gate, with mutation testing over the money-critical modules. The mutation gate itself once broke silently, reporting a perfect score while checking nothing; the fix made "not checked" a hard failure, and the triage that followed killed ten real surviving mutants in the cost-accounting code. A test suite that can catch its own tooling lying is the property that makes the other three loops trustworthy. vast.ai was chosen first because it is the cheapest market and the most hostile one, and guardrails forged against misreporting hosts and frozen docker pulls carry over unchanged to friendlier clouds. That bet needs one update: SkyPilot added a vast.ai backend in March 2026, so the marketplace is no longer exotic, and the SkyPilot-as-a-backend option in this section is now a real one rather than a hypothetical. The control-loop caution below still applies and is the reason it stays unbuilt, but the reason is now "two recovery systems fighting" and no longer "nobody else goes there". RunPod has since shipped as provider two and is exercised in the same ledger under the same rules, which is where the architecture section's seam claim stopped being a design intention. Its arrival cost three provider-specific corrections rather than an abstraction layer, and one of them generalised: the CUDA compute preflight had been probing for torch by Nebius follows for the sustained big-model work, at $2.15/hr it is the cheapest datacenter-tier H100 on-demand in our survey. This is where the vendor seam from the architecture section pays for itself: vastops will not hand-write a client per provider. The same shell-out-to-the-vendor's-tool pattern extends to SkyPilot as a provisioning backend, with one caution the design takes seriously. SkyPilot is itself an opinionated control loop, and stacking an aggressive fail-fast watchdog on top of another orchestrator's retry logic is how two recovery systems end up fighting. The likely first move is therefore a thin direct adapter held to a fixed conformance test (offer search under floors, reconcile-by-label, boot deadline, compute preflight, stop/destroy asymmetry, exit-trap, artifact hash), not a second abstraction layer. Either way the policy layer is the part that carries over untouched: hard price and reliability guardrails, the append-only cost ledger, delivery verification, the deploy gate, and the lessons loop. Bigger models need less roadmap than intuition suggests. H100 on-demand fell to $2–4/hr at specialized clouds by mid-2026, though the floor is rising again as agentic workloads soak up capacity [7], and a 70B QLoRA fine-tune plausibly fits a single 80GB card, which is a planning estimate rather than a tested recipe; the full train-checkpoint-merge-export-evaluate path gets benchmarked before that becomes a default. And one correction the VRAM arithmetic forces on the fleet thesis itself: a 70B does not fit the 24GB serving box at any useful quantization, so a 70B is never a fleet member. It is a teacher -- trained cheaply on one guardrailed rental, distilled into the specialists that actually serve. The factory that trains a 4B for a quarter can train a 70B for the price of a dinner; serving it is a separate decision the thesis does not get for free. Distributed training earns a plan rather than a build, and honesty about the boundary is the plan's first line: nothing this operation trains needs more than one node, and usually one GPU. When a job someday does, the first rung is renting a single 8-GPU box, not building a cluster -- and the second rung, NCCL across marketplace nodes, is likely a permanent buy-not-build: silent dead-worker hangs over hostile Ethernet are a research project with a maintenance tail, not an evening's guardrail. The relevant research (DiLoCo-family low-communication training [8], DeepMind's fault-tolerant asynchronous variant [9], Prime Intellect's INTELLECT-2 [10], Nous Research's DisTrO runs [11], Meta's torchft [12]) is turning training on cheap, unreliable, scattered GPUs into a supported discipline, and it stays a reading list here, activated by a real job, never by enthusiasm. The near-term backlog from the gap analysis is already half absorbed. Shipped within days of being identified: a checkpoint lifeboat (operator-side pulls on a timer plus one-command resume, chosen over instance-side R2 streaming precisely so no credentials ever land on untrusted hosts at this checkpoint size), train-versus-eval contamination checks at two tiers (lexical at build time, embedding-based on demand), the paired no-regression gate stage, and the hyperparameter sweep's decision core, with its multi-rental orchestrator deliberately waiting for the first real sweep campaign rather than shipping unexercised. Shipped since: the eval gate as a framework-level contract with an exit code, and controller-side artifact verification before destroy. Still queued: an lm-eval regression suite in the promote gate, trackio for self-hosted run history (verified importable; awaiting an on-image test rather than a claimed integration), a DPO/KTO preference-tuning path, then GRPO where answers are checkable, and multi-stage runs, where a job spans several rentals with checkpoint-resumable stages. That last one is deliberately unbuilt: the job that needs it, a prune-and-recover arm across two instance classes, is exactly the route the evidence ranks last, so it waits for a job that earns it rather than shipping unexercised. Three marketplace changes also belong on the list to absorb: vast.ai added a prepaid Reserved instance type beside on-demand and interruptible, added pagination to instance listing, which is precisely the shape of change that breaks a JSON envelope parser quietly, and listed B200/B300 at thin availability. Every one of them is the same philosophy extended: guardrails, gates, and lessons, not heroics. vastops is deliberately narrow. Today it is a two-marketplace, single-node loop, and multi-stage jobs across several rentals are named in the roadmap rather than built; the section above is a roadmap, not a claim. Serving is a home GPU box reached through an authenticated tunnel, suitable for internal tooling and product features that tolerate a home uplink, not a public low-latency SLA; the edge-native path for one product model cleared its gate on 2026-07-29 as a LoRA sibling on the provider's hosted base, and serves once the product's flag-gated router wire lands. Simulator-trained models carry documented sim-to-real caveats until a real-corpus retrain closes them. And the whole thing assumes one accountable operator: guardrails constrain automation and projects, not a team of adversarial users. The received wisdom is that custom models are for companies with ML teams and six-figure compute budgets. The counterexample here is one operator and under forty-two dollars of total marketplace spend, producing four working models that each beat a prompted model many times their size at their job, behind a process that has refused to ship five substandard candidates, watched by monitors it trained itself. The denominator belongs in the claim too: the forty-two dollars is marginal cost on top of roughly seven working days building the framework (the git history is public arithmetic), after which each additional specialist costs an evening and a quarter. That fixed-then-marginal shape is the actual thesis of the factory argument, not the headline GPU number. The pipeline is not impressive because any one part is novel. The claim is narrower and checkable: every observed failure class becomes a tested control the same day, so nothing breaks the same way twice for free -- hosts that lie get caught in minutes for cents, money leaks get closed structurally, regressions die as candidates, and the models it produces stand guard over the next run. New failure classes still arrive; the difference is what happens to them next. Most infrastructure decays. This one composts. 1. Belcak, Heinrich, Diao, Fu, Dong, Muralidharan, Lin, Molchanov (NVIDIA Research). "Small Language Models are the Future of Agentic AI." arXiv:2506.02153, June 2025. https://arxiv.org/abs/2506.02153 -- 10–30x serving-cost claim in §3.2; per-framework SLM-replaceable call estimates (MetaGPT ~60%, Open Operator ~40%, Cradle ~70%) in Appendix B. 2. Zhao et al. (Predibase). "LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report." arXiv:2405.00732, April 2024. https://arxiv.org/abs/2405.00732 3. "Cost–benefit analysis of deploying shallow, deep learning and generative models for legal text classification." 4. Meta. Llama Guard 3 8B model card, 2024. https://github.com/meta-llama/PurpleLlama/blob/main/Llama-Guard3/8B/MODEL_CARD.md -- F1 0.939 vs GPT-4 0.805 on the MLCommons-taxonomy response-classification test set. 5. LiveKit. "Solving end-of-turn detection." https://livekit.com/blog/solving-end-of-turn-detection -- model card: https://huggingface.co/livekit/turn-detector (Qwen2.5-0.5B fine-tune; ~39% false-interruption reduction, v0.4.1 vs v0.3.0). 6. Sheng et al. "S-LoRA: Serving Thousands of Concurrent LoRA Adapters." arXiv:2311.03285, November 2023. https://arxiv.org/abs/2311.03285 7. SemiAnalysis. "The Great GPU Shortage -- rental capacity," April 2026. https://newsletter.semianalysis.com/p/the-great-gpu-shortage-rental-capacity -- one-year H100 rental index ~$1.70→~$2.35/GPU-hr, October 2025→March 2026. 8. Douillard et al. (Google DeepMind). "DiLoCo: Distributed Low-Communication Training of Language Models." arXiv:2311.08105, November 2023. https://arxiv.org/abs/2311.08105 9. Google DeepMind. "Decoupled DiLoCo," April 2026. https://deepmind.google/blog/decoupled-diloco -- 88% goodput under high hardware-failure rates. 10. Prime Intellect. "INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning." arXiv:2505.07291, May 2025. https://arxiv.org/abs/2505.07291 11. Nous Research. DisTrO 15B distributed pre-training run, December 2024: https://x.com/NousResearch/status/1863622813317464157 -- hardware from Oracle, Lambda, Northern Data Group, Crusoe, and Andromeda Cluster. Psyche network (Consilience 40B): https://nousresearch.com/nous-psyche 12. Meta PyTorch. torchft -- per-step fault-tolerant training (HSDP, LocalSGD, DiLoCo, Streaming DiLoCo). https://github.com/meta-pytorch/torchft 13. "Resolution Diagnostics for Paired LLM Evaluation." arXiv:2605.30315, May 2026. https://arxiv.org/abs/2605.30315 -- 11/40 leaderboard pairwise comparisons unresolved at conventional alpha/power; common sample-size shortcuts underestimate required n by ~2x in the close-comparison regime. 14. "Small LLMs: Pruning vs. Training from Scratch." arXiv:2606.14150, June 2026. https://arxiv.org/html/2606.14150 -- at equal total token budget, from-scratch training beats coarse structured pruning of Llama-3.1-8B by 0.40 perplexity on WikiText-2; pruned initialisation wins only when the comparison is charged for the retraining budget alone. 15. Moslem et al. "Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts." arXiv:2605.28042, 2026. https://arxiv.org/html/2605.28042v1 -- expert-pruned GPT-OSS-20B at 75-87.5% compression lands competitive with, not ahead of, NLLB-200-3.3B and TranslateGemma-4B on FLORES and WMT24++. 16. "Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and Fused Chunked KL Loss." arXiv:2608.03796, August 2026. https://arxiv.org/abs/2608.03796 -- measured peak memory for an 8B teacher into a 3.2B student on one H200 at 8k context: ~103GB online dense KL, ~78GB offline, ~58GB offline with a fused chunked loss; the binding constraint is the vocabulary-sized logit tensor, not the transformer body. 17. Muralidharan et al. (NVIDIA). "Compact Language Models via Pruning and Knowledge Distillation." arXiv:2407.14679, July 2024, and the practitioner report arXiv:2408.11796. https://arxiv.org/abs/2407.14679 -- the Minitron recipe: Llama-3.1-8B to 4B recovered with 94B tokens of full distillation; the ablation retains attention intact and prunes MLP and embedding channels; width pruning wins on accuracy, depth pruning on throughput. 18. Biderman et al. "LoRA Learns Less and Forgets Less." TMLR, 2024. arXiv:2405.09673. https://arxiv.org/abs/2405.09673 -- LoRA forgets less than full fine-tuning because it learns a much lower-rank perturbation; a trade, not a safeguard. 19. Schweighofer et al. (Cognizant AI Lab). "Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies." arXiv:2605.30148, May 2026. https://arxiv.org/abs/2605.30148 -- introduces Anchored Weight Decay, a pull toward the pretraining initialisation, for Evolution Strategies fine-tuning specifically. Cited here because it is real and routinely misapplied, not because this pipeline uses it. 14. "Sub-Billion, Super-Frontier." arXiv:2606.22606, June 2026. https://arxiv.org/abs/2606.22606 -- fine-tuned Qwen2.5-0.5B at 0.83 micro-F1 on relation extraction against zero-shot GPT-5.4 at 0.69 and Claude Sonnet 4.6 at 0.66; an in-domain RoBERTa baseline also beats both, so the win is attributed to task adaptation rather than model type. Cited in the text as [14]. 15. "Small Language Models for Efficient Agentic Tool Calling." arXiv:2512.15943. https://arxiv.org/html/2512.15943 -- fine-tuned 3B variants match 7B ones; fine-tuning lifted tool-selection recall 45.0 points and invocation precision 47.9 points over base. Cited as [15]. 16. OpenAI, self-serve fine-tuning deprecation, announced 2026-05-07 (new organisations blocked immediately, existing users ended 2026-07-02, all self-serve jobs end 2027-01-06). https://developers.openai.com/api/docs/deprecations -- the stated reason is that newer base models follow instructions and formats much better and prompt-based approaches are cheaper and faster. Cited as [16]. Scope matters: this is closed-weight fine-tuning through an API, not open weights on rented hardware. 17. "The Constraint Tax." arXiv:2605.26128, May 2026. https://arxiv.org/abs/2605.26128 -- hard-schema constrained decoding on small models raises validity from 61.5% to 100% while dropping answer accuracy from 19.7% to 11.0%; Qwen2.5-1.5B falls from 91.5% to 48.0% executable accuracy on a tool-call task. Cited as [17]. 18. vLLM multi-LoRA, generally available and deployed behind AWS SageMaker and Bedrock, benchmarked with eight concurrent adapters. https://docs.vllm.ai/en/latest/features/lora/ -- cited as [18]. 19. Landscape re-check, September 2026: SkyPilot v0.13.0 (2026-07-22) and its vast.ai backend (2026-03-15, https://vast.ai/article/vast-ai-gpus-can-now-be-rentend-through-skypilot); dstack v0.20 profiles reference (https://dstack.ai/docs/reference/profiles.yml/). Both now ship hourly cost ceilings; neither ships renter-side host verification. Cited as [19]. Every number above is reproducible from this repository: eval harnesses and ship bars in -prev rollback tag. == GATE == into its log and a deploy script grepped for it, which meant the framework could not distinguish a finished run from a good one, and an eval that died before printing its number read as silence rather than as failure. Thresholds are now declared in the project's vastops.toml, the eval writes a metrics JSON, and vastops gate judges it and exits non-zero. Two behaviours are deliberate and tested: a metric the gate names but the eval never produced is a failure, not a skipped check, and a project with no gate at all exits with its own distinct code, because nothing-was-checked is not a pass. Both are the same rule applied to a gate that guards against silence as well as against bad numbers.Results
The substrate, measured
Measure Value -- - -- - Instance creates (distinct instances) 121 Create events, retries included 127 Creates that never delivered a verified working GPU 75 (62.0%) Median / p95 minutes from create to running, when a host delivered 2.4 / 8.6 Mean minutes alive per never-delivered instance ~12.3 Spend on instances that never delivered $4.96 (11.9% of total) Total marketplace spend, every failure included $41.62 Model versions gated 16 Gate refusals 6 (38%) scripts/ledger-stats.py, apart from the two gate rows, which have no derivable source: the ledger records instance lifecycle, not gate verdicts, so those two are hand-counted through 2026-09-09 and are exactly the kind of number this table exists to distrust. Two instances with no stop or destroy event are excluded from spend rather than guessed at. Two figures moved when the table stopped being typed: the earlier draft counted create events where this counts distinct instances (100 events, 95 instances, in the July window), and it included open-ended rentals in spend where this excludes them ($19.93 became $17.98 for that same window). The correction is small and it is the point: a number a command reproduces cannot drift.The models, measured
Model Job Size Baseline Fine-tune result Honest slice Train cost -- - -- - -- - -- - -- - -- - -- - ops-doctor v4 GPU-fleet action policy (STOP/INVESTIGATE/WAIT...) 8B incumbent v2: 96.7% ¹ 99.0% action acc, 1.2% missed (n=300) same-simulator eval ~$0.13 log-doctor v1 raw train.log triage (8 run states) 4B prompted 35B: 93% / 8.9% false alarm 100% acc, 0% missed, 0% false alarm, 0.56s (n=300 × 2 seeds) same-simulator; read 100% as "eval exhausted", not "model perfect" ² ~$0.35 ambie-eou v2 voice-agent turn detection 0.6B prompted 35B: 85% / 22.4% @ 524ms 98.3% acc, 0.0% early-commit, ~140ms external OOD slice: 90.0% acc, 16.4% false interruptions ~$0.30 sign-fielder v3 e-sign field auto-placement 4B prompted 35B: F1 0.421, sig recall 44.7% F1 0.922, sig recall 99.2%, valid 100% (n=86 docs × 2 seeds) advisory OOD (n=10): structural 76.1%; 51% train/val semantic-sibling rate ³ ~$3.60 incl. refused v1 The comparator question, answered
Contender In-dist F1 / recall Format-OOD F1 / recall Canaries Latency -- - -- - -- - -- - -- - Regex incumbent (train-formats only, honestly) 0.988 / 1.000 0.382 / 0.242 8/10 ~0ms Prompted qwen3.5:35b-a3b 0.597 / 0.589 0.508 / 0.483 6/10 0.8-1.7s/window ModernBERT-base fine-tune (149M, ~$0.15) 0.998 / 0.998 0.286 / 0.367 6/10 142ms/window, CPU The route question, and the experiment that settles it
Helsinki-NLP/opus-mt-es-en does Spanish to English at 75M parameters under Apache-2.0. And the money is not close: a genuine ten-billion-token recovery pass for an 8B-to-4B prune is a several-hundred-dollar rental, where specialising a 1-3B model on a hundred million tokens is a single-digit-dollar one. That is two orders of magnitude spent to arrive somewhere worse.[batch, sequence, vocabulary] logit tensor and everything the KL loss materialises from it: measured, an 8B teacher into a 3.2B student on one H200 at 8k context peaks near 103GB online with a dense KL and near 58GB with a fused chunked one [16]. Modern vocabularies make this worse rather than better, which points at the one idea the note missed entirely -- embedding plus untied lm_head on an 8B Llama is about 1.05B parameters, roughly 13% of the model, spent on tokens a single-task specialist will never emit. Trimming the vocabulary to the task's observed token set is the cheapest structured prune available, and almost nobody does it.examples/tool-caller/ is the experiment that checks it on our own bench. One task, one corpus, one evaluator, several methods: prompt the bigger model, fine-tune a small one, adapter on a shared base, quantise, and prune-then-recover, all scored by the same harness on the same splits so a difference between arms is a difference between methods. The task is schema-constrained tool calling, chosen because the output contract is rigid and the metric is exact match on a parsed call AST rather than a judge model. Everything in it is Apache-2.0 and ungated, so the matrix reproduces without a licence click.valood, which is tool-disjoint: every record offers at least one function whose schema was never shown during training, mixed in with familiar distractors, and 483 of them require calling an unseen tool. It is the direct descendant of the format-disjoint split that caught the sign-pii encoder as a surface-form memoriser, and the gap between val and valood is the number this example exists to produce. A separate canary split holds only no-call cases, where emitting any call at all is the expensive failure.What the bench actually said
timeout orphaned it. That is the single-digit-dollar figure asserted above, measured rather than estimated, with the marketplace's failure rate already inside it.arm val valood canaries valid VRAM p50 latency -- - -- - -- - -- - -- - -- - -- - sft-1.7b0.874 0.849 0.930 1.000 3.88 GB 1254 ms sft-0.6b0.864 0.814 0.908 0.998 1.63 GB 1277 ms q4-sft-1.7b0.851 0.827 0.928 0.995 1.79 GB 1638 ms base-4b prompted0.854 0.801 0.908 0.993 8.53 GB 1629 ms sft-fg270m0.788 0.726 0.853 0.953 1.11 GB 1131 ms base-0.6b prompted0.708 0.670 0.745 0.990 -- -- trim-sft-1.7b0.176 0.153 0.265 0.065 3.20 GB 12475 ms base-1.7b, base-fg270m, q4-trim-sft-1.7b) are in examples/tool-caller/results/metrics/ and none of them changes the ranking.base-1.7b was the worst zero-shot Qwen (0.653, below the 0.6B's 0.708) and sft-1.7b is the best trained arm. Selecting a base model by its off-the-shelf score would have chosen wrong.q4-sft-1.7b uses 2.2x less VRAM and is 31% slower at batch 1 (1638ms against 1254ms), because dequantisation is work paid per token with no throughput to amortise it against.sft-1.7b reached 1.0000 output validity with no constrained decoding, which is the same gap the constraint-tax literature measures from the other side.The prompt is part of the model's cost, and it is not redundant
prompt mode median tokens val valood concurrent -- - -- - -- - -- - -- - full (control) 374 0.8875 0.8500 253 signature, trained 165 0.8167 0.7917 870 names, trained 131 0.6972 0.5929 1125 signature, no retrain 165 0.8000 0.7078 -- names, no retrain 131 0.5542 0.5222 -- signature arm -- the best of them -- lands below the prompted-4B bar thesignature and 33.3 at names, andWhat this bench does NOT establish
generate() at batch 1, and decode is 97% ofsft-0.6b is 1.2GB of bf16 weights and an RTX 5090 moves aboutsft-1.7b was benched againtransformers moving 5.16docs/COMPRESSION-AND-SPECIALISATION.md carries the standing decision, including the narrow conditions under which prune-and-distil is right, which are Minitron's and not ours: you already own a large parent, you want a family of sizes, and the cost amortises over volume.The gate follows the model to the serving surface
The bet: fleets of specialists, not one know-it-all
The case against, argued properly
How the pipeline improves itself
-- gpu RTX_5090,RTX_4090,RTX_3090 emitted gpu_name=RTX_5090,RTX_4090,RTX_3090, which vast reads as one literal SKU name, so every multi-SKU pool probe the daily research protocol prescribes had been returning zero offers and reading as a thin market. -- ram 24 emitted gpu_ram>=24576, because the offer a search returns reports MB while the query it takes is in GB, so the memory floor had never matched a single host in the framework's life; live, gpu_ram>=24 matches sixty RTX 3090 offers and gpu_ram>=23552 matches none. And the spot checkpoint lifeboat, the thing that turns a dead interruptible instance into a resume instead of a loss, hardcoded vastops pull honoured the project's configured checkpoint path, so any project that wrote anywhere else had a lifeboat protecting nothing and saying nothing about it. vastops heal: the watch loop now classifies a dead session from its log tail using the same failure vocabulary the monitors use, mutates the config (halve the micro-batch and double the gradient accumulation after an OOM, preserving the effective batch size and the learning-rate schedule), re-pushes, and relaunches on the same instance, capped at two attempts. OOM is the only class that gets a remedy: NaN, like every unknown failure, stops and preserves the evidence, because a NaN has four plausible root causes and a config mutation that suppresses all of them is evidence-burying, not healing. Replayed against the morning's crash log, heal derived the exact fix the operator had applied by hand. Each applied remedy is appended to a lessons ledger with its hardware and config shape, and every launch first prints what past runs on similar shapes needed, so the same wall does not get hit twice. The second fix from the same day: a resumed instance's SSH port mapping changes and comes up slowly after a stop/start cycle, which once burned a push's retries on a dead cached port. The pipeline now waits on an endpoint-re-resolving SSH probe before it syncs.Where this goes
What it is not
Why this matters
References
Appendix: provenance of claims
examples//evaluate.py, and the raw per-arm metrics the tables are computed from in examples/*/results/metrics/, tracked on purpose, because a table whose source is untracked is an assertion rather than a result. The trained checkpoints those metrics describe are NOT in git (604 MB for the tool-caller matrix alone); run_matrix.sh regenerates them, and the ledger records what that cost. Incident and recipe history in docs/SESSION-HANDOFF.md and git log, cost events in the vastops ledger (~/.local/share/vastops/ledger.jsonl), and the monitoring ladder in docs/AGENT-LAYER.md and agents/babysitter.py.