Inference Deployment Intelligence — Target Architecture¶
Status: Frozen target architecture
Architecture version: 2.4 (2026-09-09; ADR-013 makes routed and distributed serving topology explicit without turning Apron into a scheduler or serving fabric)
Purpose: Define the complete product destination independently from delivery phases.
Governing rule¶
This document describes the product architecture, not a launch checklist. Delivery phases may change implementation order, but they do not remove or redefine architectural components. Removing, merging, or permanently narrowing a component requires an explicit architecture decision record and a new architecture version. A scope cut cannot silently become the product.
1. Model intelligence¶
Locate model artifacts through conforming sources: Hugging Face Hub first, plus local files and additional registries or object stores without changing the permanent contracts. ArtifactLocator records source kind, URI and requested revision; resolution produces a source observation with the resolved immutable revision, manifest, per-file/component content digests, sizes, license/gating observations and attributed publisher metadata. ArtifactIdentity derives from resolved contents and component structure rather than a registry namespace. The same contents at two locations retain two provenance paths but one comparable content identity; the same name with changed contents is a different artifact.
Inspect configuration, generation defaults, tokenizer/processor metadata, quantization metadata, safetensors or other weight structure, components, remote-code requirements, publisher capability claims, supported engine implementations, and existing recipes. ModelSpec is modality-open: its common core carries logical identity, resolved artifact references, license observations, exact bytes/dtype per component and lineage claims; a typed component graph carries execution mechanisms and exchange edges; CapabilitySignature separately preserves allowed input combinations, cardinalities/shapes/streaming, operation and result representation (ADR-007). A vision-language artifact can therefore contain a media encoder, projector and autoregressive decoder without being forced into one false family. Model-card fields, registry tags, provider mappings, declared lineage and evaluation results remain attributed claims. Artifact claims, pinned-engine support, configured-endpoint exposure and observed task behavior remain distinct. Activation, cache, graph and other runtime terms are fingerprinted predictions unless measured on the exact execution; unknown mechanisms or combinations remain unknown or unsupported rather than being presented as verified.
Model resolution constructs the full quantization candidate graph defined by ADR-006. It distinguishes the logical model lineage, immutable checkpoint artifacts, typed relations between artifacts, reproducible offline conversions, pinned-engine runtime transforms, and the engine/backend execution chosen for each. Native, publisher-provided, third-party, offline-derived and online-transformed candidates share one permanent contract. Specific format adapters may return unsupported; repository names and structural fingerprints never silently establish derivation, tokenizer/template compatibility, licensing or quality equivalence.
Model intelligence also discovers candidates from the accepted DecisionRequest. Exact capability signatures, license, context, tool/API support and policy facts can remove impossible candidates before paid evaluation. An independent set of modalities cannot substitute for a valid combination signature: {text,image} -> text and separate {text} -> text/{image} -> text capabilities are different. Public benchmark results and provider catalogues may supply candidate priors at their declared provenance and freshness, but they cannot establish fitness for the user's task. A managed model API is represented with its observable endpoint signature and explicit opaque fields; absence of artifact or runtime visibility is not filled with assumptions.
2. Hardware intelligence¶
Detect or describe accelerator SKU, VRAM or unified memory, device count, compute capability, interconnects, PCIe/NVLink/fabric topology, NUMA placement, CPU, host RAM, storage, operating system, driver, CUDA/ROCm/runtime versions, and container support. Convert this inventory into usable topology constraints for tensor, pipeline, data, expert, and context parallelism. Unified-memory platforms such as NVIDIA Spark are modeled explicitly rather than treated as conventional discrete GPUs.
3. Decision, task, application and serving-workload intelligence¶
The product begins with an accepted DecisionRequest, not necessarily a model name. It records the desired task or application outcome, success criteria and quality floor; serving SLO; privacy, security, license and data-residency policy; budget and time bounds; permitted providers/resources; and optimization objective. A request may constrain or preselect a model, artifact, API or target without requiring one.
TaskSuiteSpec describes what must be accomplished: versioned cases or sampling frame, input/output shapes, modalities, tools and environment, task categories, deterministic checks or rubrics, segment weights, privacy classification and representativeness claims. ApplicationSpec fingerprints the endpoint-independent system that elicits behavior: prompts and templates, agent/scaffold, tools, retrieval and other components, reasoning and sampling configuration, turn/tool/retry limits, named logical roles and invocation conditions. It does not bind roles to models. A quality result belongs to this task-and-application fingerprint, not to a model name alone.
ServingWorkloadSpec independently represents how the application must run: concurrency and request distributions, burstiness, adapter use, capacity horizon, latency/throughput/availability SLOs, and typed input/output ports. It carries the applicable distributions for tokens; image count/dimensions; video dimensions/frames/fps/duration; audio duration/sample rate/channels; document pages/components; vector dimensions/count; streaming; and operation-specific parameters such as denoising steps or guidance. Endpoint wire protocol is a separate mapping and cannot manufacture capability. Presets remain understandable, but expand into explicit accepted specifications. The compatibility WorkloadSpec envelope references both task and serving specifications. Task success is never inferred from throughput, and serving fitness is never inferred from a task score. Every measurement carries operation, direction, unit, denominator and scope so unlike signatures cannot be compared through a bare number.
4. Engine-neutral deterministic planning core¶
Use shared DecisionRequest, TaskSuiteSpec, ApplicationSpec, ServingWorkloadSpec, InferenceSolution, CapabilitySignature, ArtifactLocator, ArtifactSourceObservation, ArtifactIdentity, ModelSpec, HardwareSpec, WorkloadSpec, PlanningClaim, DeploymentPlan, EvaluationProtocol, TaskAttemptRecord, VerificationReport, EvidenceReleaseManifest, and DecisionReport contracts. Engine adapters translate exact deployable endpoint signatures and component mechanisms into engine-specific constraints and artifacts. Evaluation adapters execute replayable task protocols without owning product recommendation logic. vLLM is the initial engine implementation and resolves its generation, multimodal-input generation, pooling/scoring and speech-to-text surfaces from the pinned image rather than a hardcoded text-only family. A second LLM-serving adapter proves engine neutrality; a generated-media adapter separately proves that the architecture is not token-decoding-specific (ADR-005, ADR-007, ADR-011).
InferenceSolution is the selected object. It represents a managed model API, one self-hosted artifact/engine/target, or concrete endpoint bindings for the application's logical roles plus the resolved routing, fallback or escalation policy. It is the sole authority for those bindings; each endpoint retains its own identity, evidence and DeploymentPlan. This prevents both a single-model ceiling and duplicate routing truth.
Serving topology is typed rather than hidden behind one URL. A logical role may bind directly to one endpoint, to an application/model router, or to a replica pool; one executable endpoint may in turn be a verified sharded or disaggregated group. These are different semantics: application/model routing chooses a capability or model, replica routing chooses one interchangeable serving copy, and distributed execution makes several workers jointly serve one request. Every backing deployment preserves its artifact, engine, target, plan and evidence identity. Each attempt records the selected logical route and physical target when observable, including failover; otherwise the relevant layer is routing_opaque and evidence reuse is correspondingly limited.
Independent replicas do not pool memory. Device or node capacity may be aggregated only inside a distributed-execution group whose engine adapter verifies the exact tensor, pipeline, expert, encode/prefill/decode or other supported joint-execution mechanism and placement topology; ordinary data-parallel serving remains a replica pool. Dynamic routes are executable only when every possible data, provider/account and spend destination is inside the accepted authorization envelope; an undisclosed external destination cannot inherit permission from the public alias.
The owned planner produces candidate configurations and the claims needed to assess ADR-006's Conservative status, explains binding constraints, proposes alternatives, and distinguishes mathematical feasibility from measured performance. Conforming external planning sources—including AIConfigurator, provider recommenders and recipe planners—emit the same typed PlanningClaim: producer/version, exact inputs, proposed configuration, scope, epistemic status, provenance, calibration domain, uncertainty and unsupported or opaque fields. Every planner is a candidate generator, not qualification or selection authority. Only the common evidence-promotion and selection policy may label an endpoint Recommended or a solution Qualified, best observed or measured-efficient. AI never decides executable parameters.
Every assertion also preserves how it is known, independently of its evidence-source level. derived means a deterministic result replayable from identified inputs; proven_constraint means a scoped mathematical or pinned-source rule establishes the verdict; predicted means a named model estimates a runtime outcome and must carry provenance, calibration scope and uncertainty; measured means the exact execution produced the raw observation. A measurement from another execution—even a closely related one—may inform a prediction but is not transferred as the proposed execution's measurement.
Quantization search traverses only resolved nodes in the candidate graph. A self-hosted endpoint candidate carries separate ArtifactSpec, optional OfflineTransformSpec or RuntimeTransformSpec, ExecutionSpec, typed QuantizationSpec, compatibility evidence and evaluation-derived quality evidence. Weight, activation and KV-cache formats; dense-linear and MoE schemes; packing/scales; exclusions; calibration provenance; and selected kernels remain distinct. Memory calculations use the candidate's actual tensor/scale bytes and transformation workspace. A bootable candidate is not Recommended, and a task-capable candidate is not Qualified, until the exact solution satisfies both accepted task and serving constraints.
The planner owns the GPU-free, mechanism-aware resource calculation required before rental and download (ADR-003). It derives exact component bytes from resolved artifact manifests and dispatches calculations from the typed component graph plus accepted workload shape: MHA/GQA/MQA, MLA, sliding-window/hybrid attention, recurrent state, multimodal encoder compute/cache and media-token expansion, pooling, encoder-decoder/streaming state, or latent denoising/media decode. It never chooses a formula from text, image, audio, video, encoder, or diffusion alone. Unsupported or ambiguous mechanisms return unknown even when an engine or external planner advertises the capability. Pinned engine source supplies versioned constraints and regression fixtures; actual engine cache specifications, profile results and compatibility verdicts come only from GPU execution. Engine adapters therefore compare owned and external predictions with the instantiated engine rather than pretending that config parsing, capability discovery or a third-party recommendation is engine validation. External planning can improve or replace a covered candidate proposal; it cannot remove the transparent owned calculation/fallback path.
The planner also imports the llmcalc lineage's RunPod-tested catalogue as legacy owner-attested evidence. These records seed candidate configurations and calibration while preserving unknown historical fields; they are not silently promoted to current-version measurements. The neutral core of DeploymentPlan (model identity, hardware, workload, parallelism, precision, memory budget, expected Pareto point) is separated from an opaque engine-specific section; only the neutral core is tested for engine neutrality.
Selection and placement form one constrained decision pipeline, not an availability fallback. Cheap deterministic capability, policy and deployment-feasibility checks prune impossible solution candidates before paid evaluation. Authorization removes actions and resources outside the granted envelope. Task evaluation and serving verification reject candidates below the accepted quality or SLO constraints. The deterministic optimizer reconciles compatible planning claims, exposes disagreements and ranks the remaining solutions against the explicit objective—accepted-outcome cost, time, latency, throughput, time-to-ready, availability/resilience, privacy, or a declared multi-objective trade-off—using evidence whose epistemic status remains visible. It records the candidate set, searched and unavailable artifact/planning/provider sources, pruning reasons, evaluation coverage, ranking inputs and selected trade-off. It may call a solution measured-efficient or best observed only within a disclosed comparable set under the same task, application, serving, economics and evaluation protocols; a single qualified solution is not proven optimal, and an external source's recommended label is not a product verdict.
5. Canonical decision, solution and deployment records¶
decision-report.json is the versioned source of truth for why an InferenceSolution was selected. It contains the accepted request, task/application/serving identities, disclosed candidate set, exclusions, evaluation coverage, evidence, uncertainty, qualification state, alternatives, normalized economics and trade-offs. It references rather than duplicates immutable task attempts, verification records and endpoint plans.
deployment-plan.json remains the versioned executable source of truth for every renderer and interface. It includes the solution/endpoint identity, immutable model lineage and artifact identity where observable, typed artifact relations, any offline/runtime transform, resolved execution identity, hardware or provider target, serving workload, selected configuration, calculations, assumptions, compatibility and exact-solution quality evidence, expected performance, economics, and security risks. A single ambiguous quantization string is rejected. Compound solutions may coordinate several endpoint plans without collapsing their fingerprints.
All public schemas use semantic versions, stable identifiers, append-only evolution where possible, explicit deprecation, checked migrations, and compatibility tests. Historical plans and verification reports must remain interpretable and migratable across engine and schema generations.
6. Artifact compiler¶
Render the canonical plan into native engine commands, Docker Compose, Kubernetes resources, official or compatible Helm values, benchmark commands, CI jobs, standard inference-gateway resources and router recipes, and pluggable provider-specific deployment formats. Provider and serving-fabric renderers are optional plugins with independent ownership and freshness status; the architecture does not require permanent support for any named provider, router or scheduler.
7. Automated verification lab¶
Provision or attach hardware, or attach externally produced evidence at its declared level (ADR-002), deploy the generated artifact, observe startup, run health and functional checks, exercise the declared context/concurrency envelope, execute native benchmark sweeps, capture cold and warm behaviour, and tear down safely. Verification covers consumer, unified-memory/local, and datacenter hardware classes and continuously revalidates configurations affected by model, engine, driver, container, or schema changes.
The CPU control plane is location-neutral; maintainer-operated GPU execution runs on rented providers only (ADR-001). User-operated verification can explicitly target a compatible local, remote or rented GPU through the same ExecutionTarget contract. Ownership and GPU class do not determine evidence authority: provenance, claim scope, the execution fingerprint, attestation and independent reproduction do. Every supported GPU is an equal execution target. The scheduler chooses a run from the evidence gap, required claim scope, attributed demand, uncertainty or failure value, freshness, adequate external coverage and cost; it does not automatically include or exclude H100 or any other class. Accepted requests and local failures are direct demand; opt-in aggregate usage, multiple registries, providers, engine activity, recipes and evidence systems may add dated normalized signals. No single registry popularity counter defines eligibility or priority. Existing external evidence is ingested at its declared level and avoids duplicate execution only when it satisfies the required scope and fingerprint. Every workload measurement carries its exact execution scope, and the product refuses to present cross-fingerprint extrapolation as measurement. Hardware identity is a hard constraint when the claim itself is hardware-specific; otherwise a requested SKU may be a preference subject to explicit substitution rules.
Automated work is governed by a fixed authorization engine outside the executor's reasoning loop (ADR-010). Pluggable AuthoritySources—interactive owner approval, standing repository policy, provider/account controls and organization policy—contribute immutable constraints over the requested principal, action, resource and context. The engine combines them conservatively and returns an AuthorizationDecision with an AuthorizationEnvelope: allowed action and provider/account scopes, credential scope, permitted task-data destinations, hard target constraints, spend/runtime/teardown bounds, security/data rules and permitted adaptations. An agent may schedule and execute unattended work inside that envelope but cannot supply, widen or replace the authority that governs its own request.
A concrete target, provider or solution-member change is always reevaluated. It may proceed without new authorization when it remains feasible, stays inside the same envelope and preserves the accepted DecisionRequest; it still creates new qualification evidence when its fingerprint changes. It requires new authority when it crosses a task-data, target, budget, credential, security or adaptation boundary, and a new accepted request when it changes the task/application objective. Every authorization, scheduling and action-attempt record is durable, linked and replay-safe.
7a. Evaluation and solution qualification¶
Evaluation adapters execute EvaluationProtocols against candidate solutions and return append-only TaskAttemptRecords. Initial adapters wrap established systems such as Inspect AI for general tasks and scorers, Harbor for reproducible agentic environments, Phoenix-compatible trace experiments, external HTTP evaluation services and user-supplied commands. vLLM native benchmarks remain serving measurements and never stand in for task evaluation.
Every protocol binds the immutable task-suite, application, solution, harness, scorer/rubric, judge if any, seeds, repetitions, stopping rule, aggregation and uncertainty method. Deterministic verification is preferred where possible. Generated tasks and rubrics require separate acceptance; evaluation adapters cannot declare their own test representative. A task result is reproduced on the exact deployed solution before final qualification, so quality measured against a different API, artifact, prompt, tool graph or inference configuration cannot silently qualify the retained endpoint.
8. Evidence and compatibility database¶
Store every attempted deployment and every task attempt, including failures, retries and degraded outcomes. Deployment records preserve model commit and artifact/quantization identity; engine configuration and container digest; GPU SKU, memory and capability; selected kernel backends; runtime/library/driver versions; graph and allocator settings; parallel topology and rank; exact plan, environment and profiling/serving shape; raw non-overlapping memory observations; benchmark method; measurements, logs, timestamps, and provenance. Task attempts preserve decision, task-suite, application, evaluation-protocol and solution fingerprints; case and attempt identity; outputs and permitted artifacts; criterion scores and acceptance; traces, turns, retries and tool calls; token/cache use; time; endpoint, judge and allocated infrastructure cost; failures and raw provenance. Records declare claim_scope, production_mode, and why they were produced or ingested.
Execution requests, orchestration decisions, evidence records, proposed external actions and publication attempts have distinct identities and lifecycles. A verification result remains intact if publication is rejected or fails. Provider-account ownership and spend authorization are recorded separately from evidence authority; user-funded, maintainer-funded or provider-granted execution does not receive a stronger evidence level.
Maintainer execution has two explicit resource sources. MaintainerBaselineAllocation is the owner-controlled envelope sized from an accepted run plan and current provider prices; it must be sufficient for the Phase 1a conformance loop without promised external help. ContributedResourcePool contains provider/model-lab credits, scoped keys, quota or dedicated capacity and is additive. A contribution may expand evidence coverage or lower project out-of-pocket cost, but it cannot alter the technical candidate set, evidence promotion, user-objective ranking, publication visibility or conclusion. Resource provenance is disclosed on every affected attempt. Sponsor and partner are relationship claims used only after explicit agreement; the evidence contract itself needs neither.
A measurement is scoped to the exact fingerprint that produced it. For task outcomes this includes task suite, application, evaluator, complete solution, serving topology and observed route target in addition to the endpoint execution. A matching previous record is prior measurement evidence, not a new measurement. Cross-fingerprint observations may inform a prediction only, with source lineage, calibration scope, uncertainty and safety margin. They are never described as measurements transferred to the proposed solution. CUDA-graph bytes never cross their graph/backend/execution fingerprint as measurements; task scores never cross changed prompts, tools, routing, judges, artifacts or API identities as measurements. A measurement through an opaque alias qualifies only the observed alias behaviour under that protocol; it does not become measurement evidence for an unidentified backing deployment.
Evidence levels form a total order: maintainer-lab measured, independently reproduced, authoritative-external, community-reported, calculated-only. Within a level the newer engine version wins; across levels a disagreement is marked contested and both records are shown. Each evidence kind carries its own freshness clock (ADR-006). Every record round-trips to the ecosystem's interchange formats so that it can be consumed by, and can correct, external tools (ADR-002). Public data carries an explicit reusable data license (CDLA-Permissive-2.0 for records) and machine-readable attribution. The canonical publication unit is a portable content-addressed release whose manifest and record digests—not a host URL—define identity. Hugging Face Hub is the initial dataset mirror; conforming publishers can reproduce the same release elsewhere without changing ids, and failure or ownership change of one host does not rewrite evidence history.
9. Diagnosis and remediation¶
Accept a deployment plan plus redacted logs, tracebacks, metrics, or failed verification results. Classify model-resolution, dependency, weight-loading, quantization, memory, CUDA graph, kernel, topology, runtime KV-cache, networking, and performance failures. Return the primary cause, evidence, trade-offs, and a corrected deployment-plan.json; the corrected plan re-enters the render/deploy/verify loop.
Every remediation rule carries the failure record and corrected-boot proof that makes the mechanism mechanism_verified; without that proof it remains a hypothesis. A concrete remediation application separately records whether the accepted DecisionRequest, exact task outcome, serving workload, artifact/lineage policy and smoke semantics still pass. Only both mechanism proof and exact-solution request satisfaction produce Fixed. A healthy correction that violates an accepted task, serving or policy constraint is Alternative with trade-offs; incomplete execution is Unverified suggestion. Unknown deployment or qualification fingerprints enter the queue with trace, environment, solution/plan and accepted-request snapshot attached. Values extracted from logs are typed numerics and enums validated against trusted facts; no extracted string reaches a rendered command (ADR-005, ADR-006, ADR-011).
10. Version and compatibility intelligence¶
Inspect pinned engine containers to extract the actual CLI/API schema, defaults, supported choices, deprecations, and runtime capabilities. Diff releases, identify affected plans, run golden compatibility tests, and schedule selective revalidation. Release notes and documentation supply semantic context but are not treated as the executable source of truth.
11. Economics intelligence¶
Join task-attempt outcomes with verified serving performance, capacity and versioned provider pricing. Report cost and time per attempt; observed total attributable cost per accepted outcome; retry and failure contribution; modeled expected cost/time to acceptance; token/API charges; and self-hosted startup, transfer, idle, replica, interruption and utilization cost. Compare managed and self-hosted solutions only over a common workload horizon, quality acceptance rule and disclosed cost boundary. A more expensive token or GPU hour may produce a cheaper accepted outcome, while an inexpensive model with retries or quality failures may not. Clearly separate measured charges/performance, allocated cost, modeled extrapolation, opaque provider components and promotional pricing. For discounted or contributed execution, also separate market-equivalent price, gross attributable execution cost, subsidy/credit and project out-of-pocket cost. Ranking uses the price and entitlement available to the accepted user, not a private project grant.
12. First-class interfaces¶
Expose the same deterministic core and schemas through a web application for human discovery, CLI, structured API, local/remote MCP for agents, and repository-native CI. mondegreens/apron-action is the separately released GitHub/Marketplace adapter: it maps repository changes into canonical Apron inputs and renders versioned results as GitHub Checks, but contains no calculation, qualification, ranking, evidence-promotion or remediation authority. Interfaces cannot develop independent recommendation logic. A qualification operation evaluates authorized solution candidates against an accepted DecisionRequest and produces a DecisionReport; its CLI verb is a separate naming decision. plan performs GPU-free calculation for deployable candidates; verify performs bounded endpoint execution and exact-solution task replay, saves local reports, and tears down; deploy retains the selected managed, self-hosted or compound solution; run wraps a user-supplied execution; report only displays or exports existing records; submit separately sanitizes, previews and publishes with consent. Insufficient comparison yields a qualified or evidence-backed solution, not an optimality claim. For user-requested paid evaluation or deployment, the user supplies credential references to the user's own provider accounts and providers bill the user directly; Apron never accepts payment or resells compute. Every persistent deployment exposes continuing cost and lifecycle state.
The Action's ordinary hosted-runner path is GPU-free and labels calculations, inherited evidence, proven constraints, measurements, freshness, uncertainty and unknowns separately. A GPU job is allowed only through the same execution-target and authorization contracts as CLI or API execution; a fork pull request receives no provider credentials or side-effect authority. GitHub Marketplace, MCP registries, PyPI, GitHub Pages, portable evidence mirrors and upstream contributions distribute the same product rather than hosting divergent implementations.
13. Community verification network¶
Accept consented, redacted, reproducible success and failure reports only through an explicit submission operation separate from verification. verify keeps its report local; submit sanitizes it, shows the fields that will leave the machine, and then publishes with consent. A detected anomaly first becomes an internal AnomalyCase, not a public issue. A separate publisher first prepares the host-neutral content-addressed release or update, then applies provenance, sanitization, duplicate, confidence, destination and rate policies for each mirror, issue or pull request. Standing owner policy may authorize automatic publication; absence or rejection of that authority leaves a useful local proposal and never damages the evidence record. Use signed reports, environment attestation where available, schema validation, duplicate detection, anomaly detection, trust levels, and independent reproduction. Community evidence supplements but never impersonates automated-lab or maintainer verification (ADR-002, ADR-010).
14. Security and provenance¶
Evaluate model provenance, immutable revisions, trust-remote-code, dependency installation, container signatures/digests, licenses, exposed ports, authentication, network access, secret handling, unsafe multimodal paths, privileged deployment settings and evaluation-data disclosure. Task data, prompts, traces, outputs and evaluator inputs are private by default. Sending them to an external candidate, judge, evaluation service or provider requires the accepted data policy and explicit authorization scope. Evaluation, deployment, telemetry and publication are separate actions. Untrusted model cards, benchmark content, repository code, user logs and AI-generated explanations cannot flow directly into executable artifacts.
15. AI-native interaction with deterministic authority¶
AI interprets intent, proposes a DecisionRequest, task cases and rubrics, explains decisions, navigates alternatives, assists failure classification and may operate the lab under standing policy. It cannot accept its own proposed task definition, representativeness claim, rubric, data disclosure or optimization objective. Deterministic code resolves facts, calculates resources, executes accepted protocols, validates compatibility, aggregates disclosed measurements, filters and ranks solutions, decides whether evidence supports qualification, and enforces authorization transitions. An agent may consume granted authority but cannot grant itself more data access, spending, credentials, runtime, target adaptations or publication scope.
AI interprets and explains, deterministic code computes and generates, and untrusted inputs cannot reach executable output.
Product spine¶
Accept outcome → discover solution candidates → qualify task behavior → calculate and authorize deployment candidates → deploy → verify serving and reproduce task outcome → measure outcome economics → publish consented evidence → monitor invalidation → diagnose → correct → requalify.
The composed spine has a permanent self-hosted subflow: resolve immutable model/artifact → calculate → plan/render → GPU deploy/verify → diagnose → correct → replay the accepted request. It remains independently invocable and acceptance-tested with external evaluation services, managed inference providers, contributed resources and compound routing disabled. Task-to-solution qualification adds a decision layer around that engine; it does not replace it with a leaderboard or managed-API router.
This closed loop is the product. The resolve, resource-calculation, rendering, deployment, verification and diagnosis loop is its executable engine, not a discarded earlier product. A delivery increment may prove the loop over a narrower accepted task and candidate set, but it cannot replace the loop with a static calculator, generic leaderboard or deployment wrapper.
Architecture decision records¶
| ADR | decision | effect on this document |
|---|---|---|
| ADR-001 | Verification and computation on rented providers only | §7 lab wording |
| ADR-002 | External evidence and planning claims admitted without authority transfer; registry-neutral artifact identity; multi-source scheduling; portable evidence releases and ecosystem round-trips | §1, §4, §5, §7, §8, §11, §13 |
| ADR-003 | Apron owns the architecture-aware GPU-free predictor; GPU-executed vLLM supplies conformance evidence; llmcalc catalogue seeds legacy evidence | §1, §4, §8, §10 |
| ADR-004 | CLI/MCP direct use and repository-native checks replace generated SEO acquisition | §12 ordering |
| ADR-005 | Phase 1a/1b, evidence-obligation breadth, Phase 3 tiers, diagnosis leads, request-preserving remediation, SGLang default second engine | §9 and delivery |
| ADR-006 | Tiers, evidence order, freshness clocks, fingerprints, WorkloadSpec acceptance, two-axis remediation proof, sanitization, signing, license | §4, §8, §9, §14, §15 |
| ADR-007 | modality-open CapabilitySignature, typed component mechanisms, scoped artifact/engine/endpoint capabilities, mechanism-aware workloads; generated-media proof is distinct from second-LLM-engine proof |
§1, §3, §4 |
| ADR-008 | Upstream-first output strategy; publication begins with the first publication-eligible record, not automatically the internal fixture | §8, §13 and delivery |
| ADR-009 | Tool-neutral governance and least-privilege automation; no agent-specific scaffold is product architecture | repository policy |
| ADR-010 | Separate authority, execution, evidence and publication contracts; users authorize their own provider resources and pay providers directly | §7, §8, §12, §13, §15 |
| ADR-011 | Start from an accepted task outcome; qualify managed, self-hosted or compound InferenceSolutions through replayable task and serving evidence |
§§1, 3–8, 11–15 and product spine |
| ADR-012 | Mondegreens repository boundary; clean public Apron history; thin GitHub Action and repository-native distribution/trust surfaces | §§12–14 and delivery |
| ADR-013 | Typed logical routing, replica selection and distributed-execution topology; route attribution, no false memory pooling and fabric-neutral integration | §§3–8, 14–15 and delivery |