Per-request cost limits work when the gateway converts policy into enforceable budgets for tokens, time, memory, tools, retries, and queued work.
A family memberโs simple search can unexpectedly trigger a long model response, three retrieval passes, OCR, and several agent tools on a shared home server. Local inference has no cloud invoice per token, but it still consumes scarce accelerator time, electricity, RAM, and interactive capacity. The service needs advance estimates, live accounting, cancellation points, and clear behavior when one budget is exhausted.
Admission Estimates Reserve a Bounded Resource Envelope
Before execution, the gateway estimates input tokens, maximum output, model class, KV-cache memory, retrieval depth, tool count, and deadline. User, endpoint, and workflow policies combine into one immutable request budget that downstream services cannot silently increase.
iteration-level scheduling schedules generation at iteration granularity and dynamically batches requests with different lengths. Its design shows why actual decoding work is revealed token by token rather than known perfectly from the initial prompt. This distinction remains visible during later household testing.
Estimates should therefore reserve a ceiling without charging every request as if it reaches that ceiling. Large but low-risk background jobs can enter a queue, while interactive work may be rejected early when its worst-case memory would break an existing reservation.
Runtime Meters Enforce Tokens, Time, Memory, and Tools
Every service reports standardized usage against the request ID: prompt and generated tokens, GPU milliseconds, peak memory, CPU time, bytes read, tool calls, retries, and external operations. A central ledger subtracts usage atomically so parallel branches cannot each spend the full remaining budget.
chunked prefill scheduling studies chunked prefill and stall-free scheduling to balance throughput with decoding latency. This illustrates how one long prompt can consume service capacity in bursts unless work is divided into enforceable units. The intermediate result must remain inspectable before automation follows.
Tool budgets need semantic categories, not only counts. Ten read-only metadata calls differ from one message send or recursive file scan, so policy can limit side-effect class, target scope, output bytes, and cumulative execution time independently.
Cancellation and Partial Results Define the Budget Boundary
Cooperative cancellation propagates through retrieval, generation, and tools, and each stage checks the deadline or remaining budget before starting expensive work. Idempotency keys prevent a cancelled retry from repeating an external side effect. That boundary should be measured separately under realistic operating conditions.
fair GPU scheduling time-shares accelerator cycles to prevent request starvation and examines the cost of moving inference context. The work demonstrates that fairness controls must account for both compute time and memory state. The practical consequence appears when several sources compete for limited context.
The failure boundary is accounting without enforcement. A dashboard can report overruns while one request still monopolizes the GPU. When a limit is reached, the system must stop at a safe checkpoint, return an explicit partial result, and distinguish budget exhaustion from model or tool failure.
Test Budgets With Adversarial Request Shapes
Create requests with huge inputs, unbounded output prompts, recursive tool plans, parallel branches, retry loops, slow tools, cache misses, and cancellation during side effects. Assign distinct budgets to two users and a background service. This dependency should remain explicit in the final interface.
Use the per-user QoS boundary in per-user resource policy to record reserved and actual tokens, GPU time, peak memory, queue delay, tool operations, retry count, cancellation lag, and partial-result quality. Confirm that child spans inherit rather than reset the parent budget.
Pass only when every expensive action is attributed and no request exceeds a hard limit beyond documented cleanup work. If accurate estimation is impossible, admit conservatively and refund unused capacity instead of allowing downstream services to invent new budgets.
Tech & AI HUB
More to Read

What Features Enable a Home AI Trust Boundary Around Sensitive Files?
See how classification, capability-scoped access, isolated parsing, retrieval filters, egress policy, approvals, and audits contain sensitive home files.

What Factors Determine Whether Merkle-Tree Backups Detect Silent Change Efficiently?
Learn how chunk size, fan-out, trusted roots, cached hashes, change locality, metadata scope, and scrubbing determine Merkle backup verification cost.

What Components Enable Verifiable Backups of AI Indexes and Model State?
See how coordinated snapshots, content manifests, checksums, version locks, restore drills, and query tests prove that AI state can actually recover.

