ADR-005 — Phase 1 split into 1a/1b, Phase 3 tiered, diagnosis leads, and neutrality has two proof axes¶
Status: Accepted, 2026-09-06; amended 2026-09-07 by accepted readiness decisions D7 and D10; Phase 1a fixture/public-showcase separation accepted 2026-09-08 (D13); task-to-solution and no-permanent-engine-exclusion amendment accepted through ADR-011 on 2026-09-08; separate engine- and mechanism-neutrality obligations accepted through the ADR-007 replacement on 2026-09-08 Affects: Phase Plan Phases 1, 3, 4; Target Architecture unchanged
Context¶
Phase 1 as written in v1.3 is a team-year at the stated 8 to 20 hours per month, and the architecture freeze forbids narrowing the component set. Narrowing the evidence domain is permitted. Verification found that the only component with no prior art anywhere is log-driven diagnosis producing a corrected plan; that aiconfigurator already spans TRT-LLM, vLLM and SGLang, so engine neutrality is a baseline expectation rather than a differentiator; and that the consumer and professional hardware class has no published evidence from any lab.
Decision¶
- Phase 1a: the complete loop for one internal engineering conformance fixture on one adequate rented CUDA target: predict → render → boot → measure → record → inject failure → fingerprint → correct → reboot → replay the accepted request. The artifact is selected just before execution by recorded engineering criteria: immutable official revision, supported initial mechanism and pinned vLLM image, clear test rights, checkpoint-native fit with workload headroom, reproducibility, isolation of the product path, availability and cost. Popularity and showcase value are not acceptance criteria. Qwen3-8B BF16 and RTX 4090 are viable candidates, not hardcoded requirements. Exit: internal evidence record #1 and remediation record #1 exist and round-trip to the three interop formats.
- Record #1 is engineering evidence, not the project's public demonstration. Publishing it is optional and requires the normal publication gates. A later public showcase is selected independently from current demand or a real failure and demonstrates the differentiated end-to-end loop—prediction beside measurement, diagnosis, request-preserving correction and reproducibility—not merely that a recent model can boot.
- Phase 1b: breadth is defined by evidence obligations, not a fixed GPU-SKU matrix. The cohort must exercise small, mid and large models; consumer single-GPU, professional/datacenter and multi-GPU execution; architecture and quantization diversity; prediction-error cases; and the six principal injected-failure classes: out-of-memory, engine core initialization failed, max-model-len exceeds derived, dtype incompatibility (Gemma bf16), tensor-parallel divisibility, quantization versus compute capability. The evidence-gap scheduler selects the exact supported
ExecutionTargetfor each obligation under ADR-002. - Diagnosis is the lead component of Phase 1. A rule is
mechanism_verifiedonly when a matching record proves that its correction removes the identified failure and restores engine health; otherwise it is ahypothesis. Applying that rule to a request produces a separate request-satisfaction result against the accepted workload, artifact/lineage policy, smoke semantics, and applicable performance and quality constraints. Only an application with both mechanism and request-satisfaction proof is presented asFixed. A healthy deployment that violates an accepted constraint isAlternative with trade-offs; an unexecuted or incompletely evaluated correction isUnverified suggestion. Unknown fingerprints enter the queue with the raw trace, environment and plan attached. - Phase 3 tiers: Tier 0 is GPU-free per-tag schema/constraint extraction and Apron prediction recomputation, explicitly without vLLM validation. Tier 1 produces GPU-executed evidence on any supported target selected by evidence gap, claim scope, demand, uncertainty/failure value, freshness, adequate external coverage and cost. Tier 2 ingests external evidence at its declared level and may satisfy a Tier 1 evidence need only when its claim scope and execution fingerprint are adequate. Budget figures are scoped planning estimates under INV-16; the scheduler and INV-14 caps, rather than a hardware-class exclusion, bound actual spend.
- Engine neutrality and inference-mechanism neutrality are separate Phase 4 obligations. A second LLM-serving adapter defaults to SGLang for datacenter parity when no accepted decision demonstrates a better target; it proves that the neutral core is not vLLM-specific. A generated-media adapter over a pinned Diffusers-based service, Triton deployment or better evidence-selected target proves that the same contracts are not token-decoding-specific. Completing either adapter cannot satisfy the other obligation. Their execution order follows accepted demand, evidence value and authorized cost, but both use the permanent schemas defined before Phase 1 records exist (ADR-007).
- llama.cpp, MLX, CPU engines and other runtimes are not permanently excluded: built-in auto-fit reduces the value of duplicating one capability, but a real task/solution comparison, hardware target or evidence gap may justify an adapter. Existing comparable serving evidence is consumed when adequate; exact-task cross-engine evidence is produced when the accepted
DecisionRequestrequires it (ADR-011). - The workload measurement rule stands: every workload measurement carries a hardware-class and capability-signature scope; the product refuses cross-class or cross-signature performance extrapolation.
Consequences¶
- Phase 1 exit gate is reachable solo; the externally dependent clause ("at least one real failure has completed the loop") is satisfied by injected failures in 1b and by external failures thereafter, and the plan says which.
- Phase 3 recurring cost is bounded and stated.
- Conservative / Recommended / Candidate are defined in ADR-006 so that Phase 1 can emit them honestly.
- The product cannot claim modality neutrality after adding only another token-serving engine, or engine neutrality after adding only a media pipeline.