Skip to content

Inference Deployment Intelligence — One-Page Phase Plan

Plan version: 2.4 (2026-09-09; applies ADR-001 to ADR-013, including routed/distributed serving topology) Architecture dependency: Target Architecture 2.4. Delivery rule: Each phase adds operational depth or coverage. It does not redefine the destination.

Product and distribution thesis: Every accepted task protocol, evaluated solution, resolved model, diagnosed failure and verified deployment becomes product evidence that other tools can consume. The product joins task outcome, serving behavior and exact deployment rather than competing with evaluation harnesses or managed inference providers. Distribution is manufactured by closing that evidence loop and emitting artifacts in ecosystem shapes (ADR-002, ADR-011); it is not assumed to be free, automatic or independent of product quality. CLI and MCP are the primary direct-use surfaces; the Apron GitHub Action brings the same deterministic core to model, engine, deployment and application repository changes (ADR-004, ADR-012).

Execution environment: the GPU-free control plane may run on a user or maintainer machine, in CI, or as a hosted CPU service. Maintainer-operated GPU work runs on rented providers; a user may explicitly select a compatible local, remote or rented GPU. vLLM compatibility and boot evidence come only from execution inside the official pinned image with detected target hardware. ExecutionTarget isolates provisioning/transport from the stable verification contract (ADR-001, ADR-003).

Phase 0 — Contracts and truth model

Start condition for schema implementation: Accept or explicitly defer the review findings that change schema meaning, and pin representative source artifacts for every external format being mapped. Each source entry records its URL or repository, immutable revision, retrieval date, digest, license, and exact schema/example/validator or consumer code used. If no formal schema exists, that absence is explicit; pinned fixtures replace an imagined contract.

Two implementation decisions are start conditions rather than later refinements, because both are consumed by the frozen contracts and cannot be changed afterwards without invalidating existing identity or fingerprints (ADR-006, INV-42, INV-43). The record-serialization and digest mapping — RFC 8785 canonical form, SHA-256, multihash prefix, with published cross-implementation test vectors — must exist before a record schema is written, because INV-35 derives release identity from it. The determinism ports for wall-clock time, identifier generation and randomness must exist before record writers and orchestration, because every fingerprint and every deterministic claim depends on there being no ambient source. Neither is an architectural component; both are defined in ADR-006 and expressed in engineering-standards.md §4–5.

This is not a general project-entry checklist. Schemas are versioned and migratable from their first draft; there is no blanket provisional label, arbitrary record-count trigger, maintainer-conversation gate or publication-organization gate on beginning them. Provider credentials, gated-token handling, external spend limits, timeout and teardown become mandatory before the first provider run. Code/data licenses, sanitization permissions, and publication location become mandatory before the first publication. A vLLM maintainer conversation and the eleven named items in the concept assessment §10 are non-blocking research actions; only concrete evidence that invalidates an accepted contract or the Phase 1a loop can stop or revise the affected work.

Define the permanent automation boundaries in the first contracts: accepted DecisionRequest; side-effect-specific ActionRequest; AuthorityContribution, AuthorizationDecision, AuthorizationEnvelope, OrchestrationDecision, ActionAttempt with evaluation/execution/publication specializations; verification and task-attempt records; AnomalyCase; ProposedExternalAction; provider-account ownership, spend authorization and data-destination permission. They have separate identities and lifecycles. Pluggable authority sources contribute versioned constraints; a fixed deterministic authorization engine combines them and neither scheduler, evaluator nor executor can widen task-data access, target, credential, budget, runtime, adaptation or publication scope. Restart and duplicate-event fixtures prove that a paid evaluation/provider run or external mutation cannot occur twice (ADR-010, ADR-011).

Objective: Establish the stable conceptual spine before interfaces proliferate.

Define versioned DecisionRequest, TaskSuiteSpec, ApplicationSpec, ServingWorkloadSpec, compatibility WorkloadSpec, InferenceSolution, CapabilitySignature, ArtifactLocator, ArtifactSourceObservation, ArtifactIdentity, ModelSpec, HardwareSpec, PlanningClaim, EvaluationProtocol, TaskAttemptRecord, endpoint DeploymentPlan, VerificationReport, EvidenceReleaseManifest, and DecisionReport schemas (ADR-002, ADR-007, ADR-011). Preserve the neutral-core / engine-specific split, evidence order and conflict rule, attempt outcomes, provenance, risk, economics, migration rules and adapter contracts. Resolve registry locations to immutable revisions and content/component digests before calculation or execution. Every assertion independently declares its epistemic status: replayable derived, scoped proven_constraint, uncertainty-bearing predicted, or exact-fingerprint measured. Cross-fingerprint observations may calibrate a prediction but cannot be transferred as a measurement. Owned and external planners emit candidate claims through one contract; only the common qualification/selection policy emits a verdict. Evidence releases are content-addressed and reconstructable independently of any mirror.

Task semantics and serving load are separate. Adopt the field's serving vocabulary verbatim: ISL/OSL, TTFT, TPOT/ITL, perceived versus total TPS, P50/P90/P95/P99, inference-only versus end-to-end scope, and goodput at SLO. Task protocols separately preserve cases/sampling frame, application/agent/tool graph, deterministic checks or rubrics, judge identity, seeds, repetitions, stopping rules, aggregation and uncertainty. Encode qualification as a dependency graph: candidate discovered → capability eligible → solution identity resolved; then task and applicable deployment/serving evidence may be scheduled in the cheapest authorized order; verified retained endpoint + task evaluated + serving SLO verified + task outcome reproduced on retained solution → solution qualified. Managed APIs use provider_opaque for hidden runtime fields rather than simulated self-hosting evidence.

Build golden fixtures for one single self-hosted solution, one managed API solution and one compound planner/executor/vision or fallback graph. The single and compound solutions pass through the same decision/evidence APIs while every endpoint retains its own identity and plan. Evaluation-adapter fixtures cover a deterministic scorer, a model judge, a failed/retried attempt and a private task rejected from an unauthorized external destination. Capability fixtures cover text-to-text; text-plus-image-to-text with a forbidden-combination counterexample; file and realtime audio transcription; embedding and query/document scoring; text-to-generated-image; audio-to-text-to-audio composition; a checkpoint ability not exposed by its engine; and an engine-resolved ability whose calculator remains unknown. Component/mechanism fixtures still cover dense, MoE, quantized, long-context, multimodal, hybrid-attention, multi-GPU and unified-memory execution. Apply ADR-006, ADR-007 and ADR-011 fingerprint, acceptance, freshness, privacy and outcome-economics rules.

Define selection and placement as a permanent deterministic obligation graph: cheap capability/policy/deployment-feasibility pruning and authorization establish eligible resolved solutions; accepted task evaluation and applicable deployment/serving verification may then be scheduled in the cheapest authorized order; exact retained-solution task reproduction plus every applicable serving SLO is required before objective-based ranking can promote a qualified solution. Implementations may interleave measurements to minimize information cost but cannot skip a qualification obligation. The request distinguishes hard constraints, preferences and permitted substitutions. DecisionReport preserves every considered solution, rejection reason, evidence state, evaluation coverage and trade-off. Golden tests cover hardware-specific measurement, general deployment substitution, a high-throughput task failure, a high-quality serving-SLO failure and an incomparable managed-versus-self-hosted cost boundary.

Apply ADR-006's execution fingerprint as the measurement boundary. VerificationReport stores initial total/free/requested memory, model/weight memory, persistent consumption, transient peak headroom, non-PyTorch increase, CUDA-graph estimate/applied/actual values, available KV-cache memory, safety buffer and the exact profiling shape as separate fields. It must not normalize vLLM's composite peak_activation_memory into a pure activation measurement or add graph memory twice. Cross-fingerprint observations enter the planner only as attributed predictions with uncertainty.

Implement ADR-006's complete quantization candidate graph in the initial contracts: ModelLineage, ArtifactSpec, ArtifactRelation, OfflineTransformSpec, RuntimeTransformSpec, ExecutionSpec, typed QuantizationSpec, evaluation-derived QualityEvidence, and CompatibilityEvidence. QualityEvidence references the task attempts and protocol that produced it; it is not a free-standing assertion. Quantization fixtures remain unchanged in breadth, and no format may require changing the public schemas or APIs.

Define the calculator contract before freezing its outputs: Apron owns GPU-free prediction and dispatches by typed component execution mechanism plus accepted workload shape, never by modality or a broad architecture-family label. The mechanism registry represents ordinary MHA/GQA/MQA, MLA, sliding/local/hybrid attention, recurrent/Mamba state, media encoder/processor/cache and media-token expansion, pooling, encoder-decoder/streaming state, latent denoising and media decode. A mechanism branch is executable only when conformance-tested; all others return unknown without hiding an engine-resolved capability. It consumes immutable artifact/component metadata, exact checkpoint bytes where available, versioned engine constraints and qualified calibration records. The pinned GPU path records instantiated layers, processors, backend cache/state specs and profiled memory as conformance and measurement, never as a CPU-only promise.

Import the 101 unique non-Gemma llmcalc catalogue entries as owner_attested_legacy_boot. Preserve original empirical VRAM, configuration and provider fields; keep absent artifact revision, engine/image digest, exact hardware, workload, log and date explicitly unknown. The import seeds candidates and replay selection but does not satisfy current verification gates. Do not carry forward llmcalc's generic KV, activation, MoE or quantization fallback formulas.

ModelSpec retains a common identity core plus a typed component/mechanism graph; CapabilitySignature separately records operation, allowed input combinations and output representation (ADR-007). ADR-011 separates task semantics into TaskSuiteSpec and traffic/SLO semantics into ServingWorkloadSpec. The initial vLLM adapter is not text-hardcoded: it extracts runner type, supported tasks, implementation, multimodal combinations and limits, and exposed endpoints from the pinned image for generation, multimodal-input generation, pooling/scoring and speech-to-text. Discovery and each Apron stage remain separate support states, so schema breadth never fabricates calculator or runtime evidence.

Define MaintainerBaselineAllocation and ContributedResourcePool as distinct resource sources before any provider backend executes. Build a budget proof for the accepted Phase 1a run plan from current provider quotes, expected transfer/storage, every planned attempt, correction, serving benchmark, task replay and teardown. The baseline envelope is owner-controlled and cannot be set by the scheduler. Contributed provider/model-lab capacity has provenance and may reduce project out-of-pocket cost or expand coverage, but it cannot change evidence authority, selection or user-facing economics (ADR-010, INV-29).

Exit gate: The same fixtures round-trip through versioned schemas and migrations without losing meaning; generated endpoint plans and decision reports are deterministic; unsupported, opaque, unknown and estimated states are explicit. Every DeploymentPlan fixture still exports losslessly for shared fields to the three pinned ecosystem shapes (ADR-002). Single, managed and compound InferenceSolutions share one API; changed members, routing or modality transitions change the solution fingerprint. Direct-endpoint, identical-replica, heterogeneous-alias, failover, opaque-router and verified-sharded fixtures preserve their distinct topology semantics and route attribution; replica capacity cannot masquerade as sharded capacity, and every possible dynamic destination is checked against authorization. Evaluation adapters cannot leak private fixtures, transfer a task score across fingerprints or hide failed attempts. Every capability fixture preserves combination semantics, result shape and its separate artifact/engine/endpoint/evidence state. The generated-image fixture resolves a real immutable pipeline into text-encoder, denoiser and decoder mechanisms without claiming a vLLM endpoint or executable media proof. The accepted Phase 1a run plan is budget-feasible with ContributedResourcePool = 0.

The exit gate also requires that every quantization fixture resolves through the same candidate API; artifact, transform and execution identities survive round-trip; unsupported adapters fail explicitly; name similarity cannot promote lineage; actual tensor and auxiliary-scale bytes feed memory calculation; no endpoint reaches Recommended without accepted task and serving evidence; and no solution reaches Qualified without exact-solution task reproduction plus every applicable SLO.

Phase 1 — First complete vLLM product loop

Objective: Prove the executable deployment/diagnosis engine and then the complete task-to-qualified-solution promise over a bounded, credible evidence domain (ADR-005, ADR-011).

Phase 1a — one internal conformance fixture, one rented CUDA target, the whole loop. Immediately before execution, select the artifact and adequate target by a recorded engineering rule: immutable official revision; clear test rights; support in the pinned vLLM image and the implemented memory mechanism; checkpoint-native fit with headroom for the accepted workload; reproducibility; isolation of the product path; availability and total expected workload cost. Hourly SKU price alone is not the objective. The fixture is not selected for fashion, launch messaging or provider promotion. Qwen3-8B BF16 on RTX 4090 is a viable fallback combination, not a hardcoded requirement.

Execute the fixture against the owner-controlled baseline envelope even if provider participation is zero. If an eligible contributed resource is available, the scheduler may use it only under the same target-selection rule and must retain the counterfactual market-equivalent cost; the contribution cannot determine the model, target, result language or publication. Failure of the baseline budget proof is a plan/envelope defect to resolve before execution, not a reason to assume future sponsorship.

A GPU-free Apron prediction resolves the complete candidate graph for the selected lineage using immutable artifact metadata, typed artifact/transform/execution identities, the implemented mechanism-aware calculation branch, version-pinned engine constraints, applicable llmcalc legacy evidence, and the declared target. Unsupported candidate adapters remain visible as unsupported; they do not collapse the graph or fabricate a flag-based artifact. Actual vLLM compatibility, instantiated-layer/backend cache spec, boot, kernel, profile and memory verdicts execute in the pinned vllm/vllm-openai:<tag> image with the rented GPU visible and become measurements.

The fixture has a real accepted DecisionRequest, small deterministic TaskSuiteSpec, fingerprinted ApplicationSpec, explicit ServingWorkloadSpec and replayable EvaluationProtocol. The task suite exists to prove the contract and request-preserving correction, not to claim that the engineering fixture is a market-representative benchmark. A conforming evaluation adapter records every TaskAttemptRecord before and after deployment correction. The fixture also exercises artifact graph resolution; hardware detection; actual tensor/scale byte calculation; inspectable cache/non-KV/budget prediction; evidence-bound result states; deployment-plan.json; native vllm serve and Docker Compose renderers; the adapted llmcalc provider flow; serving measurement via vllm bench; prediction deltas; and one deliberate, source-backed, reversible failure injection whose correction still satisfies the accepted task and serving request. A hardware-specific incident or task score is not reused across a non-matching fingerprint. CLI and local MCP expose every step, including the run wrapper.

CLI execution contract is a Phase 1a deliverable. The qualification operation executes an accepted protocol for the constrained fixture and produces a DecisionReport; its CLI verb is not frozen by ADR-011. Broad candidate discovery is not claimed by that one run. apron plan produces GPU-free endpoint candidates against the accepted request. apron verify PLAN --target TARGET runs bounded deployment verification and exact-solution task replay, saves local verification and task-attempt records, and tears down. apron deploy DECISION_OR_PLAN --target TARGET_OR_ENVELOPE retains an explicitly managed solution. apron run -- ... wraps an existing execution for observation and diagnosis. Paid evaluation, judge and provider actions show the applicable data destination, maximum estimated cost and hard deadline and require confirmation or standing authorization. apron report RECORD_ID is read-only display/export. apron submit RECORD_ID is a separate opt-in operation that sanitizes, previews and publishes. Evaluation, execution and submission never imply one another.

Exit gate 1a: internal deployment evidence record #1, task-attempt evidence and remediation record #1 exist. They carry exact decision/task/application/evaluator/solution and execution fingerprints, provider, instance, digest, predicted-versus-measured memory, serving observations, mechanism outcome and accepted-request outcome. The corrected endpoint reproduces the accepted deterministic task result; this proves the machinery, not general model quality. The public verify command runs through the selected rented target and produces locally saved schema-valid records that pass sanitization and provenance validation. A no-GPU target returns a typed non-measurement outcome; target loss and every failed run still trigger teardown. report is read-only and submit requires explicit consent. The permanent self-hosted acceptance suite passes with external evaluation services, managed inference providers, compound routing and contributed resources disabled (INV-30). Publication and provider participation are not exit conditions.

Public showcase is a separate product/distribution artifact. It is selected from a real accepted use case and demonstrates the expanded sequence: task definition and acceptance → disclosed managed/self-hosted or compound candidate set → task evaluation → prediction beside deployment measurement → serving-SLO check → exact-solution task reproduction → outcome economics → diagnosis or regression correction → requalification. It cannot be merely a recent model boot or general leaderboard. It follows ordinary data, publication and provider-brand rules and does not retroactively determine the engineering fixture.

Phase 1b — breadth by evidence obligation. Build a cohort that exercises task-quality and serving trade-offs; at least one managed versus self-hosted comparison with a normalized cost boundary; one single-endpoint and one compound-solution fixture; small, mid and large models; consumer single-GPU, professional/datacenter and multi-GPU execution; architecture and quantization diversity; prediction-error cases; and the six principal deployment failure classes. Store every evaluation failure, retry and deployment failure as a first-class attempt. The evidence-gap scheduler chooses candidates and exact supported resources by expected decision value inside authorization; no provider, model or GPU class is privileged or forbidden by name.

Publish mondegreens/apron-action when the structured Apron input/output schema is stable for its first repository check. The Action pins a released core, performs GPU-free resolution/calculation/compatibility analysis, and renders a Check Summary that preserves evidence state, coverage, freshness, uncertainty, invalidation and unknowns. One external fixture repository must consume an immutable Action release and reproduce the expected result. Provider or local GPU dispatch is admitted only after that path passes the normal execution-target and authorization conformance suites; it is not simulated in a hosted CPU runner.

Diagnosis is the lead component and a first-class product surface. Its executable correction is produced by versioned deterministic rules from parsed error values (typed numerics and enums only), immutable model metadata, hardware facts, compatibility records, and scoped proving records. A rule becomes mechanism_verified only when corrected boot restores engine health; otherwise it remains a hypothesis. Each application then re-runs the accepted request and records request_outcome independently. The public result is Fixed only when both the mechanism and request are verified, Alternative with trade-offs when the engine is healthy but an accepted constraint is violated, and Unverified suggestion when proof is incomplete. An unknown failure remains explicitly unrecognized and enters the internal queue with trace, environment, plan, and accepted-request snapshot attached.

An AI agent can propose and test diagnosis rules autonomously because it has access to every input the analysis requires: the engine source code at the pinned tag (the exact line where the error fires and the condition that triggers it), the model metadata (safetensors index, config.json), the hardware spec (compute capability, VRAM, supported dtypes), and the full error trace. The agent reads the engine's source to identify the binding constraint — not guessing a fix but tracing why the condition failed and selecting a correction from the engine's own valid options (--help=all output, dtype compatibility tables in source). INV-6 applies: values must be typed numerics or enums, not interpolated strings. A proposed rule starts at status: hypothesis; a matching corrected boot may promote the rule to mechanism_verified, but the application is not Fixed until the accepted request is re-run successfully (INV-2). A wrong analysis or an objective-breaking correction produces a failed/violated remediation record, not a verified fix. The agent reads the engine's source to identify the binding constraint.

Publish exact environment records, measurements, evidence badges, independent freshness clocks, raw methodology, externally sourced authoritative evidence with attribution, and the prospective rate at which monitored solutions satisfy states in the qualification graph. A metric such as "boot-verified within 48 hours" must publish its numerator, complete eligible denominator, observation period, model/hardware cohort, multi-source eligibility rule, source contribution and exclusions. Generated model and error web pages are not a Phase 1 deliverable and are not scheduled (ADR-004); portable content-addressed evidence releases, their initial Hugging Face dataset mirror, the MCP server and the repository-native Action are the public surfaces.

Exit gate 1b: A user or agent can complete the entire loop from accepted task outcome to qualified retained solution, bounded DecisionReport, or honest no-qualifying-candidate result through CLI and MCP. A separate fixture repository consumes a pinned mondegreens/apron-action release and receives an equivalent GPU-free result and evidence-state Check Summary from the same core. Task and serving failures remain independent; exact-solution reproduction is required. Every one of the six deployment failure classes still completes fingerprint → typed extraction → deterministic correction → deployment → mechanism proof → accepted-request re-execution without case-specific code. Outcome economics retains failed attempts and retries. Single and compound solutions use the same APIs; managed and self-hosted candidates compare only under normalized boundaries. No component depends on manually copied recommendation logic.

Phase 2 — Production and agent interfaces

Objective: Put the proven core inside real engineering workflows.

Add structured API, remote MCP, conforming Inspect AI and Harbor evaluation adapters, trace/dataset import through a Phoenix-compatible adapter, and an external HTTP/user-command adapter. Extend the released Apron Action from its GPU-free repository check to any required authorized verification workflow only through the same execution-target contracts. Add Helm values rendering, security and provenance reporting, schema migration tooling, consented submit, signed community reports, outcome economics using versioned price data, and the thin web decision view over the same core. MCP, web, CLI, Action and CI consume identical decision, task, solution, plan and report contracts; no surface may merge evaluation, deployment, telemetry or submission authority.

Bring-your-own-provider evaluation, verification and deployment. A user explicitly supplies credential references for the user's own model, judge, evaluation and compute providers. apron qualify, verify or deploy previews data destinations, authorization envelope and objective; executes only permitted task data and resources; enforces hard cost/runtime/lifecycle limits; and records why candidates and targets were selected. An unavailable preference may be substituted inside the envelope; a task-data, hardware-specific or other hard-boundary violation cannot. Providers bill the user directly. Provider-account ownership does not change evidence authority, and execution never implies telemetry or publication.

Apron does not collect money, resell compute, settle provider charges or operate pass-through billing. It receives only the explicit authority needed to access resources in the user's account. Provider credentials come from the approved environment/keychain or delegated secret mechanism, never CLI flags, plans, reports or stored logs (ADR-010, INV-13).

Exit gate: An agent and a CI pipeline can independently accept or reference a decision request, run a supported task protocol, resolve and plan solution endpoints, validate, render, diagnose and retrieve evidence without a human translating web output. Historical Phase 1 records remain readable. On-demand evaluation and verification use the same canonical pipelines as scheduled work and preserve their distinct evidence scopes.

Phase 3 — Continuous compatibility lab

Objective: Make freshness and accumulated evidence the defensible asset at solo-sustainable cost (ADR-002, ADR-005).

Three tiers. Tier 0, execution-free and rewindable: detect material changes in engine schemas/constraints, model/API identity, application or task-suite revisions, evaluator/judge protocols, provider behavior/pricing and serving workloads; recompute derived facts and invalidate affected qualifications without manufacturing runtime evidence. Tier 1, produced evidence: execute the minimum authorized task, endpoint or serving experiments whose information can change a decision, selected by evidence gap, claim scope, demand, uncertainty, unresolved-failure value, freshness and cost. Tier 2, ingested evidence: ingest external deployment, serving and evaluation evidence at its declared scope and level. External evidence prevents duplicate execution only when it covers the required task/application/solution/protocol or deployment fingerprint. A public benchmark cannot qualify the user's task; a boot cannot become performance evidence; a different endpoint's score cannot qualify the retained solution. Recently changed models, APIs and task/application revisions enter the queue by material impact rather than popularity alone. Failure-taxonomy expansion, anomaly detection, adapters and public exports continue.

Agentic maintenance pipeline (2026-09-07)

Evidence: solo planning tools die when the biweekly cadence becomes a human treadmill (TheBloke: manual-repetitive at accelerating scale, stopped after 1 year). Study of 115,466 repos (arxiv 2507.21678): maintainer activity interval is the #1 predictor of cessation. An automated pipeline keeps the interval at zero without human effort.

The Tier 0 → Tier 1 → diagnosis → upstream-preparation pipeline runs as a durable agentic workflow, not a human checklist. The maintainer defines or approves standing policy and monitors exceptions rather than executing the normal cycle. Each job records its stable id, deduplication key, policy version, state transitions, inputs and outputs. The executor can act unattended inside the policy but cannot modify the budget, credentials, deadlines, destinations or action classes that govern it (ADR-010).

Release detection has two layers, reactive and anticipatory:

  1. Reactive (cron): A scheduled GitHub Action polls the vLLM container registry for new tags. Detects the release when it drops. Simple, reliable.
  2. Anticipatory (signal intake, agentic): The agent watches upstream activity BEFORE the tag drops — merged PRs, RFCs, release branches, issue clusters (Phase 3 signal intake). When a merged PR changes a versioned compatibility constraint or kernel selector, the agent identifies the exact artifact/backend/hardware records that may be affected and queues them for GPU testing; it does not generalize one scheme's rule to all FP8. When the cron detects the tag, the agent already knows what to test — no cold-start analysis on release day.
step automation does authority/publication boundary
Release detection Registry events or bounded polling detect tags; signal intake watches upstream PRs/RFCs/issues and pre-computes scoped impact. Read-only sources and polling cadence are declared by policy.
Tier 0: tag diff, schema/constraint extraction, prediction recompute Runs GPU-free for every detected tag and produces prediction-only changes. Standing repository policy may authorize commit, push and merge after required checks; the job cannot change that policy.
Tier 1 dispatch Ranks evidence gaps, filters for feasibility and authorization, then selects an exact target against the recorded evidence objective. Separate authorization fixes the allowed envelope, provider credentials, hard target constraints, substitutions, cost/runtime bounds and teardown before dispatch; the selected SKU need not be fixed when it is only a preference.
Tier 1 review Creates or updates a deduplicated internal AnomalyCase for an unknown fingerprint or significant prediction delta. No issue is created merely because an anomaly was detected.
New diagnosis rule Reads pinned source, proposes a typed hypothesis, executes an authorized correction run, and records mechanism and accepted-request outcomes separately. Promotion remains deterministic and evidence-backed; external publication is a separate action.
Upstream contribution Renders a verified record into the upstream shape and creates a ProposedExternalAction. A separate publisher submits, drafts or retains it locally according to destination-specific standing policy, provenance, deduplication, confidence and rate controls.

The normal cycle can run unattended when standing policy authorizes it. A human is needed when policy demands review, authority expires, a new action class is requested, or an exception cannot be resolved. Changing a review requirement changes policy; it does not change evidence, plan or report schemas.

Budget-aware Tier 1 scheduling (2026-09-07)

Principle: INV-14 is an independent authorization boundary, not a value owned by the scheduler. Budget-aware scheduling optimizes evidence value inside a granted envelope; provider/account controls and backend-enforced per-job limits, runtime and teardown bound the consequence if scheduling or reasoning fails. Availability may remove a candidate but cannot by itself make another candidate efficient.

The scheduler receives, but cannot edit, the remaining authorized allocation, its validity window, allowed targets, hard constraints, preferences, substitution rules and action classes. It combines those facts with pending releases, the ranked evidence-gap queue, provider availability and estimated candidate cost to maximize weighted evidence value rather than raw record count. Changing the concrete target reruns feasibility, authorization and objective ranking; it requests new authority only when the replacement falls outside the existing envelope, and a new accepted request when it changes the claim or deployment objective. A candidate already covered by adequate current external evidence is skipped unless its recorded reason identifies a different scope, fingerprint, contested result or reproduction need. Cost estimates guide selection but never expand authorization; numerical defaults remain explicitly scoped estimates under INV-16.

The baseline and contributed pools remain separately visible to this scheduler. Provider/model-lab support is admitted only as credits, scoped credentials, quota or capacity under ADR-010; it never becomes cash revenue, recommendation authority or an assumed future allocation. Scheduling may consider the real marginal project cost of a contributed run when choosing which evidence to acquire, but the resulting DecisionReport ranks solutions using the cost entitlement available to the accepted user and exposes market-equivalent price, subsidy and project out-of-pocket cost separately.

When the authorized GPU allocation is reduced, the pipeline runs fewer or cheaper evidence jobs under the same contracts. When it is zero, Tier 0 and local prediction can still run, while new paid execution waits in a typed blocked_by_authority state. Expired credentials, unavailable targets or publication denial are likewise explicit states; automation never treats them as permission to widen access.

GPU dollar protection (2026-09-07)

Source: advocacy grilling — "what if the whole budget is consumed by model downloads on spinning resources or repeated boot failures?" This is a real failure mode observed in practice.

Four rules that prevent GPU budget waste, each enforced before GPU time is billed:

  1. Never download on GPU-billed time. Model weights are pre-staged to a provider persistent volume using a CPU-only instance (\(0.01-0.10/hr). The GPU instance starts with the volume already attached. Download time (30-60 min for large models) is billed at CPU rate, not GPU rate. For a 65 GB model: ~\)0.05 on CPU vs ~$2.00 on a 4090. The provider script provisions the volume, stages the weights, THEN starts the GPU instance.

  2. Never boot a candidate that the static prediction proves impossible. The CPU prediction may reject a candidate from model bytes, declared VRAM, structural constraints, or a version-pinned incompatibility. A passing prediction is permission to measure, not vLLM validation; unknowns remain explicit and GPU execution is still authoritative.

  3. Do not reproduce a known failure without an evidence reason. Before dispatching a Tier 1 run, the agent checks the diagnosis rule table. If a matching rule is mechanism_verified, it may skip intentionally recreating the original failure and execute the corrected plan directly, but it must still replay the current accepted request before returning Fixed. If the rule is a hypothesis, the corrected plan is an Unverified suggestion until execution establishes the mechanism and request outcomes. Reproducing the original failure remains eligible only when the fingerprint/scope is contested, stale, incomplete, or itself the evidence gap.

  4. Estimate cost before each dispatch. The agent computes: model download size (from safetensors index), estimated boot+benchmark time (from prior records for similar model sizes), GPU hourly rate (from provider API), and total estimated cost. If estimated cost exceeds remaining budget allocation, the model is skipped and the next one in the ranked list is tried. The cost estimate is logged beside the actual cost in every record, calibrating future estimates.

These four rules mean: GPU dollars are spent only on feasible combinations that haven't already been tested, with weights already on disk, and with a cost estimate checked against the budget before the instance starts.

Demand-driven model ranking (2026-09-07)

Source: advocacy grilling session — "when users submit their task, the prompt is analyzed agentically." Principle: the agentic pipeline should spend GPU budget on what people actually deploy, not just what's popular on HuggingFace.

Every CLI invocation (apron plan, apron deploy, apron run, apron verify, read-only apron report, and explicit apron submit), MCP call and repository-Action check may produce the categorical local signals defined by the telemetry contract. Without telemetry consent, an Action result remains in the invoking repository workflow and sends no usage event to Mondegreens. No prompts, filesystem contents, credentials, raw reports or measurements are transmitted without the applicable consent. Verification and telemetry consent remain separate; submit governs evidence publication.

These signals feed the Tier 1 ranked list alongside HuggingFace 30-day downloads:

signal what it tells the ranking fallback without users
Model requested via CLI/MCP Which models are actually being deployed, not just downloaded HF downloads (includes CI, researchers, non-deployment)
Accepted task/application class (coding, chat, RAG, batch) What people need the solution to accomplish — prioritizes task suites and application adapters Maintainer-defined public task fixtures
Required capability signature and modality combination Whether demand is text, vision, audio, video, embedding/scoring or generated media, without collapsing distinct input combinations Published task fixtures segmented by signature
Hardware in requests What GPUs people actually have — prioritize catalogue columns Assumed hardware from rented SKUs
Concurrency and context length Whether "fits" is meaningful — fits at concurrency=1 ≠ fits at concurrency=6 Default ServingWorkloadSpec
Outcome (plan worked / failure class / corrected) Closes the prediction-vs-outcome loop from real deployments, not just lab boots Lab boots only

With 0 users, the ranking falls back to HF downloads and maintainer-defined presets — exactly the current design. With users, the ranking shifts to actual demand. The product gets smarter about WHERE to spend its budget without additional maintainer effort.

Calibration loop. A user's local verify result enters the shared evidence corpus only through explicit submit. Accepted community records contribute prediction-versus-measured deltas across hardware beyond the maintainer-rented lab while retaining their evidence level and execution fingerprint. Local reports that are not submitted remain local and never silently calibrate the public system.

Exit gate: A new vLLM release triggers Tier 0 automatically and Tier 1 selectively within a fixed external authorization, identifies stale plans, and produces updated evidence without a manual rewrite. Execution remains inside cost/runtime/target limits and converges on teardown after cancellation, crash, timeout and target loss. Replayed events create neither duplicate paid runs nor duplicate publications. Demand signals affect ranking only under their consent contract. The normal authorized cycle completes across consecutive releases; blocked actions remain durable and explain which authority or evidence is missing.

Phase 4 — Prove engine, mechanism and accelerator/backend neutrality independently

Objective: Demonstrate independently that the system is not vLLM-specific, not token-decoding-specific and not coupled to one accelerator/runtime backend.

Carry three non-substitutable adapter obligations. The second LLM-serving engine defaults to SGLang unless demonstrated demand identifies a better engine-neutrality target. The generated-media adapter uses a pinned Diffusers-based service, Triton deployment or better evidence-selected engine and proves a different component graph, serving payload, unit system and resource mechanism. The accelerator/backend proof selects one practical non-CUDA or otherwise independent execution backend from available owned, rented or contributed hardware and current engine support; Apple Silicon/MLX, CPU, ROCm, XPU or another backend is eligible, but none is promised before its adapter and evidence exist. One implementation may satisfy more than one obligation only when each conformance record remains separately attributable. For each, implement the resolver, capability signatures, prediction constraints, renderer, hardware-executed verifier, diagnostics, protocol mapping and version extraction through the existing contracts. External serving comparisons may be consumed at their scope; task qualification across engines is produced when required by an accepted decision because only the exact application/task protocol can establish it. Extend through the frozen solution and evidence contracts; do not fork the neutral core.

Exit gate: One conformance record proves the same exact capability signature through vLLM and the selected second LLM-serving engine. A separate conformance record proves a generated-media signature and mechanism graph through the selected media engine. A third record proves an exact supported signature on the selected independent accelerator/backend and leaves unsupported operations explicit. Accepted decision requests produce task, serving, endpoint and economics evidence with signature-appropriate units without changing interface contracts or corrupting historical records. The neutral InferenceSolution, endpoint DeploymentPlan and every Phase 1 record remain valid unchanged across all three proofs.

Phase 5 — Standard and sustainable service

Objective: Turn the evidence network into durable shared infrastructure.

Expand independent task and deployment reproductions, evaluation/trace integrations, hardware and inference-provider adapters, private-model and organizational policy support, continuous qualification/fleet regression monitoring, and reproducibility attestations. Keep public calculations, schemas, methodology and consented reusable evidence open under their declared licenses. This phase is reached when external tools consume the decision/evidence schemas or qualification interfaces and real applications use requalification to govern changes. Income is not a project goal; sustainability means authorized evaluation and GPU operating cost is covered without Apron becoming a payment intermediary.

Exit gate: External tools depend on the schemas or evidence dataset, verification coverage grows beyond maintainer-rented hardware, and operating cost is supported without putting basic planning behind a paywall.

First 90 days — External validation protocol

The technical Phase 1 exit gate and the market checkpoint are separate. Phase 1 asks whether the system works. This checkpoint asks whether real use causes evidence to create further use.

Before launch, instrument privacy-respecting events for CLI, MCP and repository-Action invocations and publish the counting rules. Over the first 90 days, measure these predeclared targets without substituting page views, Marketplace impressions, workflow references, badge views or search impressions for product use:

  • At least three independent external failures, submitted by at least two external users, complete fingerprint → diagnosis → modified plan → successful-boot evidence without custom logic for each case.
  • At least ten plans are generated for model IDs that were not manually seeded by the maintainer.
  • At least five generated artifacts are used through a measurable product action (rendered and deployed via CLI or MCP, or copied via an instrumented action).
  • At least three diagnoses cause a user or agent to generate or execute a modified plan.
  • At least one signed or reproducible verification report arrives from hardware outside the maintainer-rented lab.
  • At least one user or agent returns and performs another planning, diagnosis, or verification action.
  • At least one evidence record receives an unsolicited external reference in a GitHub or Hugging Face issue, community discussion, provider document, or is consumed by an external tool.
  • At least one repository outside the Mondegreens-owned conformance fixtures completes a pinned Apron Action check on a material model, engine, artifact, workload or deployment change and completes another check after a later material change.

Base-rate prognosis: low initial distribution remains the explicit estimate, not a product risk disguised as certainty. The preserved comparison shows that planning tools generally attracted less visible adoption than a reusable catalogue of validated consumer-GPU Compose files, but volatile star totals are not copied into the phase plan or treated as capability evidence. The checkpoint targets above are hypotheses, and a miss is read against their published counting rules; the oracle needs a consumer, not popularity. Consequence adopted: verified records are also rendered into a recipes-style catalogue of per-model, per-card Compose files and serve commands with the measured numbers beside them, generated from records by the existing renderers (ADR-004 amendment 2).

Decision rule: The distribution flywheel is supported only if every product-pull target above is met. If either the repeated external remediation loop or the unsolicited-reference target fails, the flywheel is not demonstrated. Other partial results are reported as mixed rather than rounded up to success. A failed checkpoint does not silently shrink the target architecture; it triggers an explicit decision to revise positioning, acquisition surfaces, phase split, or continued founder investment.

Phase governance

The public claim is: new well-formed checkpoints receive a calculated plan automatically; selective testing promotes them; unknowns enter a queue. Any stronger automation percentage is published only after prospective measurement from the complete attempt database. Public evidence language must distinguish calculated, inherited, boot-verified, workload-verified, and remediation-verified states, and every workload number carries its hardware-class scope. Progress is measured by successful plans, completed verification attempts, corrected failures, repeated use, artifact use, external reports, unsolicited references, schema and dataset consumers, time-to-promotion, and revalidation coverage, not page views.

Operating capacity assumes the community contributes no dependable reduction in maintainer work. The agentic pipeline compresses recurring compatibility work into an automated cycle governed by an external standing policy: AI can trace pinned source, propose and test corrections, and format upstream-ready actions, while deterministic gates govern evidence promotion and a separate publisher governs external mutation. The maintainer's recurring work is policy and exception review, not execution of a checklist. Capacity is bounded by authorized GPU spend, provider availability, download/boot wall-clock and upstream cadence. Numerical operating assumptions remain explicitly labeled estimates with their scope under INV-16. Phase boundaries can move. The frozen target architecture cannot be narrowed through this plan; any proposed removal requires an explicit architecture decision and version change.