Small-Model Training Substrate on Rented GPUs

Marketplace GPUs are cheap and hostile. We make them dependable.

An RTX 3090 rents for eleven cents an hour, pricing a 4B fine-tune at twenty-five cents. But the hosts misreport hardware, freeze mid-pull, and keep billing after training dies. tuneharness wraps marketplace compute with post-boot bf16 preflights, 30-minute idle watchdogs, append-only ledgers, and deterministic deploy gates that refuse substandard models.

Total Compute Spend
$41.62
Across 14 active training days with every boot failure and dead host included.
Delivery Failure Rate
62.0%
75 of 121 distinct instance creations died before training began. Absorbed without data loss.
Spend on Dead Hosts
$4.96
Fail-fast detection killed bad rentals in minutes, keeping total dead-host waste under $5.
Gate Refusals
6 of 16
Refusals are the point: candidate regressions die in staging before ever touching live serving.

Interactive CLI Walkthrough

A stdlib-only policy layer that verifies everything

Explore real terminal sessions recorded from actual marketplace runs. Click each workflow tab to inspect how the harness provisions, preflights, judges, and heals training jobs.

tuneharness cli session
LIVE REPLAY

Defensive Engineering

Four invariants that turn strangers' rigs into reliable compute

Cloud orchestrators assume friendly datacenter hosts. Marketplace hosts are consumer rigs run by strangers. tuneharness assumes every host is lying until verified by renter-side code.

INVARIANT 01

Post-Boot CUDA Preflight

An instance that passes nvidia-smi can still fail on real compute due to driver/toolkit mismatches. The harness runs an immediate torch.cuda.init() and bf16 matmul. Bad hosts are terminated and replaced in ten seconds, preventing the 105-minute silent idle burns typical of naive rentals.

Time to detection: 9.4s median
INVARIANT 02

Provider-Neutral Policy Seam

Exactly one module talks to vendor CLIs (vast.ai and RunPod). The policy layer (guardrails, quarantine, reputation scoring, cost ledger, watchdog, and deploy gates) never mentions the provider. Swapping marketplaces changes one adapter while all hard-won rules carry over untouched.

Conforms across vast.ai & RunPod
INVARIANT 03

Money as a First-Class Failure Mode

The append-only ledger records create, stop, and destroy events with timestamps, rates, and reasons. Automated watchdogs sample GPU utilization on a cron and stop idle instances. Stop and destroy are asymmetric: automation may stop, but only the operator destroys.

Total bad-host loss: $4.96 across 75 failures
INVARIANT 04

Deploy Gate with Teeth

No model reaches production without beating its prompted generalist baseline on held-out data. Evaluations are judged on the worst of two seeds, tested on 10 exact failure canaries, and evaluated on format-disjoint out-of-distribution slices. Six candidate models have been refused to date.

Exit code contract: missing metrics = failure

Empirical Economics

Interactive GPU Marketplace Risk & Cost Calculator

Calculate the true cost of small-model training. Compare managed hyperscaler clouds against naive marketplace renting (with frozen pulls and silent idle burns) versus the guarded tuneharness pipeline.

Training Runs Per Month12 runs
Training Hours Per Run3.0 hrs
Marketplace Host Failure Rate62%
Historical measured rate from our 121-instance ledger: 62.0%
Hyperscaler Cloud Cost:$64.80
Naive Marketplace (w/ Unnoticed Waste):$49.25
tuneharness Guarded Marketplace:$29.53
Idle-Burn Waste Prevented:$19.72
Net Operator Savings:$35.27

Deterministic Deploy Gate

Where regressions die as candidates

A deploy gate that only ever passes models is theater. This gate earns its keep by saying no. Inspect real candidate promotion trials from our production history.

Active Candidate:
Final Gate Verdict

Empirical Grounding

The models, measured on held-out tasks

All evaluations are held-out and judged on the worst of two seeds. Baseline comparison is qwen3.5:35b-a3b (35B MoE generalist) running on our 24GB serving box.

ModelTaskSizeBaseline ReferenceFine-Tune ResultHonest Slice / CaveatTrain Cost
ops-doctor v4GPU-fleet action policy8BIncumbent v2: 96.7%99.0% action acc, 1.2% missedSame-simulator eval~$0.13
log-doctor v1Raw train.log triage (8 states)4BPrompted 35B: 93% acc100% acc, 0% missed, 0.56sRead 100% as eval exhausted~$0.35
ambie-eou v2Voice turn detection0.6BPrompted 35B: 85% @ 524ms98.3% acc, 0.0% early-commit, 140msExternal OOD: 90.0% acc~$0.30
sign-fielder v3E-sign field placement4BPrompted 35B: F1 0.421F1 0.922, sig recall 99.2%Advisory OOD structural: 76.1%~$3.60
tool-caller sftSchema-constrained tool calling1.7BPrompted 4B: val 0.854val 0.874, valood 0.8491.0000 valid with zero grammar tax~$0.26

Continuous Hardening

The Four Operational Feedback Loops

The interesting property is not any single feature. It is the loop: every operational failure becomes code the same day, and the pipeline gets harder to break every time it breaks.

LOOP 01

Incidents to Code

An operational lesson may not live in human memory or documentation alone; it must land as a commit making the defect structurally impossible. Includes WAN flap grace windows, SSH proxy fallbacks, and autonomous OOM healing.

Remedies recorded in immutable lesson ledger
LOOP 02

Protocolized Research

Training days start with automated knowledge checks for new architectures, toolchain updates, and live market rates. Research triggers action only if a candidate plausibly beats the incumbent by 2 points or halves cost.

Avoids unexercised architecture churn
LOOP 03

Pipeline Models Babysit

The always-on monitoring ladder runs deterministic detection first, small specialist triage second (ops-doctor and log-doctor), and frontier LLM escalation only for ambiguous residue with capped spend budgets.

Autonomous stop requires deterministic proof
LOOP 04

Tests Police the Tests

The harness maintains 100% branch coverage with mutation testing over money-critical paths. When a mutation gate broke silently, the fix made unverified execution a fatal exit, killing ten surviving mutants in cost accounting.

Mutation-tested cost and lease accounting