Skip to content

ADR-010 — Automation authority, spending and publication are separate contracts

Status: Accepted 2026-09-08 by D2; evaluation/data-destination and inference-solution amendment accepted through ADR-011 on 2026-09-08; provider-resource independence and cold-start amendment owner-approved 2026-09-08
Affects: target architecture §§4, 7, 8, 11–15; framework invariants and extension points; Phases 0–3 External controls checked: NIST SP 800-162 ABAC, OASIS XACML 3.0, Cedar authorization, GitHub deployment environments, REST API best practices, REST API secondary/content-creation limits

Context

The compatibility lab must be highly automated to remain sustainable for one maintainer. The earlier phase plan nevertheless gave one agent a compound role: rank evidence gaps, spend against a budget, execute GPU jobs, judge anomalies, promote diagnosis work, and create external issues or pull requests. A cap passed as an argument to that same executor is not an independent authority boundary, and an observed anomaly is not automatically fit for publication.

The solution is not to remove autonomy or require a human confirmation for every operation. It is to make authority, execution, evidence and publication distinct, versioned contracts. Standing owner policy can then safely authorize unattended work without allowing an executor to enlarge its own powers.

Authority must also be separated from placement quality. Permission to use several targets does not mean that the scheduler may deploy to whichever one happens to be available. The accepted deployment objective still has to be satisfied, and the selected target must be justified against the user's optimization objective. Conversely, authorizing a class of acceptable substitutions should not force a new confirmation merely because the initially preferred SKU is unavailable.

Inference providers, GPU platforms and model labs may have a legitimate reason to contribute credits, scoped credentials, quota or dedicated capacity: independently produced records can demonstrate their systems, improve integrations and bring qualified workloads. That support is valuable only if it cannot buy a conclusion, suppress an unfavorable record, distort user economics or become a prerequisite for the lab to start and continue its minimum evidence cycle.

Decision

  1. The permanent contracts are separate. The truth model distinguishes an accepted DecisionRequest; side-effect-specific ActionRequest; orchestration/scheduling decision; authorization decision and envelope; evaluation, execution and publication attempts; task, verification and other evidence records; proposed external action; publication attempt; and immutable authority-source references. Provider-account ownership, spend authorization and task-data disclosure permission are separate from evidence authority.
  2. Authority sources contribute policy; a fixed authorization engine decides. Interactive owner approval, standing repository policy, provider/account controls and organization policy are typed AuthoritySource implementations. They contribute versioned constraints over principal, action, resource and context. The deterministic core authorization engine combines those contributions; deny wins, indeterminate or absent explicit permission cannot authorize a side effect, constraints intersect, and an unknown required obligation denies execution. The executor cannot act as an authority source for its own request.
  3. Authorization returns an envelope, not an arbitrary target or model choice. The envelope fixes permitted action classes; candidate, model, judge, evaluation and compute providers/accounts; credential scopes; task-data destinations; hard target constraints; maximum spend; runtime/lifecycle/teardown; security/data rules; publication scope; and allowed adaptation rules. It may preserve user preferences without turning them into hard constraints. The decision is bound to the ActionRequest, accepted DecisionRequest, context and source versions.
  4. Qualification and optimization remain independent of authorization. Cheap capability, policy and deployment-feasibility rules remove proven-invalid InferenceSolution candidates; authorization removes prohibited actions and resources; accepted task and serving evidence qualify candidates; the deterministic optimizer ranks comparable survivors against accepted-outcome cost/time, latency, throughput, time-to-ready, resilience, privacy or a declared balance. Availability is an input, never the objective by itself. The decision record preserves candidates, exclusions, evidence coverage and trade-offs.
  5. Substitution is reevaluated, not automatically denied or accepted. A changed target, provider or solution member may proceed only when it remains feasible and qualified, stays inside the same authorization and data-policy envelope, and preserves the accepted objective. A new authorization is required when it crosses a hard constraint, task-data destination, cost/security/credential scope or adaptation rule; a changed task/application objective requires a new accepted request. Any material solution change requires new qualification evidence even when separately authorized.
  6. Controls exist outside the reasoning loop. Backends enforce idempotent teardown and job deadlines. Provider/account quotas, protected environments, narrowly scoped credentials or an equivalent external control bound the maximum consequence when supported. Estimated cost guides scheduling; it never replaces a hard limit.
  7. Orchestration is durable and replay-safe. Jobs and external actions have stable identifiers, deduplication keys, state transitions, retry/backoff rules and audit records. Restarting the controller must not duplicate a paid evaluation/judge call, provider run, dataset publication, issue, pull request or deployment.
  8. Detection is not publication. A failure or prediction anomaly first becomes an internal AnomalyCase. A separate publisher evaluates provenance, sanitization, duplicate state, destination policy, confidence and rate limits. Standing policy may authorize automatic Git operations, issue creation or pull-request submission; otherwise the action remains a local proposal or draft. The executor cannot grant itself publication permission.
  9. Verification does not imply sharing. verify saves locally; submit is the explicit user-consented evidence-publication path. Maintainer-lab automation may publish only under its separately configured publisher policy.
  10. Apron is never the payment intermediary. For user-requested model, judge, evaluation or compute-provider work, the user explicitly supplies credential references for the user's own provider accounts and those providers bill the user directly. Apron uses granted credentials only for authorized actions and records account ownership, authorization and attributable provider cost without changing evidence authority. Credentials remain in approved secret sources and are neither stored in plans/reports nor accepted as payment. Maintainer work uses separately controlled maintainer or provider-granted resources. Apron does not collect money, resell services or compute, settle charges or implement pass-through billing.
  11. Claims preserve how they are known. A result assertion is derived when it follows deterministically from identified inputs, proven_constraint when a scoped mathematical or pinned-source rule establishes it, predicted when it models a runtime outcome, and measured only when the exact execution produced the observation. Predictions carry provenance, calibration scope and uncertainty. Prior or cross-fingerprint measurements may inform a prediction but are never transferred as the proposed execution's measurement.
  12. Efficiency claims are bounded. best or optimal is never a free-standing label. The product may report a discovered or task-evaluated candidate, calculated deployment candidate, verified endpoint, qualified solution or measured-efficient solution. Measured-efficient means best observed among a disclosed comparable solution set for the same accepted task, application, serving, economics and evaluation protocols inside the authorization envelope. One qualified solution is not proven optimal.
  13. Contributed resources never govern conclusions. A provider or model lab may contribute account credits, a scoped key, quota or dedicated execution capacity; Apron does not accept cash or resell that capacity. The contribution is a resource-ownership and authorization fact, not evidence authority, candidate preference, endorsement or permission to review, delay, edit or suppress a result. The scheduler continues to choose experiments by evidence value inside policy; the optimizer continues to rank solutions by the accepted user objective. No contribution may require exclusivity, favorable placement or a predetermined result. Public language uses sponsor or partner only after an explicit agreement and applicable brand approval.
  14. Cold start is independently fundable. MaintainerBaselineAllocation and ContributedResourcePool are separate authorization/accounting sources. The accepted Phase 1a execution plan, including failed attempt, corrected boot, serving measurement and task replay, must fit a maintainer-funded or otherwise unconditionally controlled baseline envelope when contributed resources equal zero. The owner sets that hard envelope from current provider quotes and the exact run plan before execution; the scheduler cannot set or enlarge it. Contributions expand coverage or reduce project out-of-pocket cost but cannot be assumed by an exit gate, reduce evidence requirements or become the only implementation path.
  15. Subsidy does not rewrite economics. Every contributed or discounted run separately records public/list or otherwise reproducible market-equivalent price, gross attributable execution cost, subsidy/credit applied and project out-of-pocket cost. User-facing solution comparison uses the cost actually available to the relevant user under the accepted request; project-only grants do not make a candidate appear free or more efficient. Unknown price components remain unknown.

Required failure tests

  • The executor cannot raise a budget, widen a credential scope or extend its own deadline.
  • A concrete target change inside the authorized substitution rules and hard limits proceeds without new authority; the same change outside any hard boundary is denied and records which boundary requires new authorization.
  • A hardware-specific verification cannot substitute a different GPU, while a GPU preference in a deployment request can be substituted only when the accepted workload and optimization objective remain satisfied.
  • The optimizer cannot select a merely available candidate over a demonstrably better permitted candidate without recording an objective- and evidence-based reason.
  • Cross-fingerprint evidence cannot be emitted as a measurement, and a single measured candidate cannot be labeled optimal.
  • An evaluation adapter cannot send private task, prompt, trace, output or rubric data to an external candidate or judge absent a permitted destination in both the accepted request and authorization envelope.
  • Replaying a job cannot duplicate a paid task attempt or judge call, and failed/retried attempts cannot disappear from outcome economics.
  • Cancellation, timeout, controller crash and target loss all converge on idempotent teardown.
  • Replayed events cannot duplicate paid executions or external publications.
  • An anomaly remains useful and queryable without creating a public issue.
  • A publication rejected by policy leaves the evidence record intact and records the rejection reason.
  • Changing from manual to standing autonomous authorization changes policy, not plan/report schemas.
  • With ContributedResourcePool = 0, the authorized Phase 1a plan remains executable inside MaintainerBaselineAllocation; if it does not fit, the plan or envelope must be explicitly revised before execution rather than assuming future credits.
  • Granting a provider credit cannot change candidate order, evidence level, publication language or result visibility when all technical and user-objective inputs are unchanged.
  • A fully subsidized run reports both the subsidy and the non-zero market-equivalent/user-relevant cost; it cannot be ranked as free unless the accepted user has the same durable entitlement.

Consequences

  • The full autonomous lab remains part of the destination architecture.
  • Human review is a configurable policy decision, not a permanent bottleneck or an implicit absence.
  • Authorization, feasibility and placement optimization can evolve independently without weakening one another.
  • Provider availability causes a new deterministic selection inside the existing envelope, not an unprincipled "deploy to whatever is available" fallback.
  • Git automation remains allowed, while external side effects are scoped, auditable and revocable.
  • Provider-granted project resources fit the same account-ownership and target contract without becoming evidence endorsements, recommendation authority or a user subsidy unless that user has the same explicit entitlement.
  • A useful provider relationship can grow from reproducible records, integration quality and qualified demand without making the provider load-bearing or compromising an unfavorable result.
  • Cold-start proof and the minimum evidence loop have an owner-controlled funding path; contributed capacity is additive rather than existential.
  • The core can be implemented and tested without inventing a billing platform or coupling itself to a particular agent harness.