Observe / Replay / Calibrate¶
This guide covers the end-to-end pipeline for validating BLIS simulator accuracy against real inference servers: observe real latencies, replay the captured trace through the DES, and calibrate by comparing results.
# Quick example: observe a real server, replay through the simulator, compare
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--workload-spec workload.yaml --trace-header trace.yaml --trace-data trace.csv
./blis replay --trace-header trace.yaml --trace-data trace.csv \
--model qwen/qwen3-14b --results-path results.json
./blis calibrate --trace-header trace.yaml --trace-data trace.csv \
--sim-results results.json --report calibration.json
Pipeline Overview¶
The observe/replay/calibrate pipeline has three stages:
| Stage | Command | Input | Output |
|---|---|---|---|
| Observe | blis observe |
Workload spec or distribution params + real server | TraceV2 (header YAML + data CSV) |
| Replay | blis replay |
TraceV2 files + simulator config | SimResult JSON |
| Calibrate | blis calibrate |
TraceV2 + SimResult JSON | Calibration report JSON |
Data flow:
WorkloadSpec YAML ──► blis observe ──► TraceV2 (header.yaml + data.csv)
│ │
▼ ▼
Real Server blis replay ──► results.json
│
▼
TraceV2 + results.json
│
▼
blis calibrate ──► calibration.json
Why three separate commands? Each stage is independently useful. You can observe without replaying (to collect latency baselines), replay without calibrating (to test simulator behavior on real traces), or re-calibrate with different simulator configs without re-observing.
blis observe¶
Dispatches requests to a real inference server, records per-request timing (TTFT, E2E latency, token counts), and exports the results as a TraceV2 file pair.
Required Flags¶
| Flag | Type | Default | Description |
|---|---|---|---|
--server-url |
string |
"" |
Inference server URL |
--model |
string |
"" |
Model name for API requests |
--trace-header |
string |
"" |
Output path for TraceV2 header YAML |
--trace-data |
string |
"" |
Output path for TraceV2 data CSV |
Workload Input (one required)¶
Four input modes are available. At least one must be provided per invocation:
| Mode | Flags | Description |
|---|---|---|
| Named preset | --workload <name> --rate <N> |
Standard workload from defaults.yaml; identical token distributions to blis run --workload <name> |
| Workload spec | --workload-spec <file> |
Multi-client workload from a YAML file |
| Distribution synthesis | --rate <N> |
Single-client workload with custom token distributions (see Distribution Synthesis Flags) |
| Closed-loop | --concurrency <N> |
Fixed pool of virtual users; arrival is response-driven (token distributions from Distribution Synthesis Flags) |
Flag reference:
| Flag | Type | Default | Description |
|---|---|---|---|
--workload |
string |
"" |
Preset name (chatbot, summarization, contentgen, multidoc); requires --rate |
--workload-spec |
string |
"" |
Path to WorkloadSpec YAML (alternative to --workload or --rate) |
--rate |
float64 |
0 |
Requests per second; required for --workload preset mode and distribution synthesis |
--concurrency |
int |
0 |
Number of closed-loop virtual users; mutually exclusive with --rate |
Combinations that produce an error:
| Combination | Error |
|---|---|
--workload without --rate |
preset requires a rate |
--workload + --workload-spec |
mutually exclusive |
--workload + --concurrency |
mutually exclusive |
--rate + --concurrency |
mutually exclusive |
--workload-spec + --concurrency |
use clients[].concurrency in the spec file instead |
Note
--workload-spec takes priority over --rate if both are provided — the spec is used and --rate is ignored. All other distribution synthesis flags (--prompt-tokens, etc.) are similarly ignored when --workload-spec is active.
Optional Flags¶
| Flag | Type | Default | Description |
|---|---|---|---|
--api-key |
string |
"" |
Bearer token for server authentication |
--server-type |
string |
"vllm" |
Server type (vllm, tgi, etc.) |
--max-concurrency |
int |
256 |
Maximum simultaneous in-flight requests |
--warmup-requests |
int |
0 |
Number of initial requests to exclude from trace |
--prewarm-duration |
duration |
0 |
System priming phase before real workload (e.g., 60s). Sends small fixed requests at low concurrency to warm CUDA/EPP/memory. 0 = disabled |
--no-streaming |
bool |
false |
Disable streaming (use non-streaming HTTP) |
--seed |
int64 |
42 |
RNG seed for workload generation |
--horizon |
int64 |
0 |
Observation horizon in microseconds (0 = from spec or unlimited) |
--num-requests |
int |
0 |
Maximum requests to generate (0 = from spec or unlimited) |
--lazy-generation |
bool |
false |
Alpha (#1441/#1443): stream requests from the generator instead of pre-generating the full slice (same flag/semantics as blis run). Supports every workload class — multi-session reasoning (SingleSession=false, #1458), concurrency clients (concurrency > 0, #1459), and time-varying / per-window workloads (#1460); there is no eager fallback. Default (off) dispatch behavior is unchanged |
--think-time-ms |
int |
0 |
Think time in ms between response and next request (concurrency mode only) |
--api-format |
string |
"completions" |
API format: completions or chat |
--unconstrained-output |
bool |
false |
Do not set max_tokens (let server decide output length) |
--min-tokens |
int |
0 |
Set min_tokens in request body; requests server to generate at least N tokens before EOS. Set equal to --output-tokens for exact output length control (0 = omit). Compatible with --unconstrained-output: min_tokens is still sent, max_tokens is still omitted |
--timeout |
int |
300 |
HTTP request timeout in seconds (per request); increase for slow servers or large-prefill workloads |
--rtt-ms |
float64 |
0 |
Measured network round-trip time in milliseconds |
--defaults-filepath |
string |
"defaults.yaml" |
Path to defaults.yaml containing preset definitions (preset mode only) |
--record-itl |
bool |
false |
Record per-chunk timestamps for ITL calibration (forces streaming per request; mutually exclusive with --no-streaming; use with --itl-output) |
--itl-output |
string |
"" |
Output path for ITL CSV file (default: <trace-data>.itl.csv when --record-itl is set) |
Distribution Synthesis Flags¶
Used when --rate or --concurrency mode is active (ignored when --workload-spec or --workload <preset> is provided):
| Flag | Type | Default | Description |
|---|---|---|---|
--prompt-tokens |
int |
512 |
Average prompt token count |
--prompt-tokens-stdev |
int |
256 |
Prompt token standard deviation |
--prompt-tokens-min |
int |
2 |
Minimum prompt tokens |
--prompt-tokens-max |
int |
7000 |
Maximum prompt tokens |
--output-tokens |
int |
512 |
Average output token count |
--output-tokens-stdev |
int |
256 |
Output token standard deviation |
--output-tokens-min |
int |
2 |
Minimum output tokens |
--output-tokens-max |
int |
7000 |
Maximum output tokens |
--prefix-tokens |
int |
0 |
Shared prefix token count |
Examples¶
Named preset mode — drive the server with a standard workload (same shape as blis run --workload chatbot):
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--workload chatbot --rate 10 --num-requests 100 \
--trace-header trace.yaml --trace-data trace.csv
Workload-spec mode — multi-client workload from a YAML spec:
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--workload-spec workload.yaml \
--trace-header trace.yaml --trace-data trace.csv
Rate mode — quick experiment with distribution synthesis:
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--rate 10 --num-requests 100 \
--prompt-tokens 256 --output-tokens 128 \
--trace-header trace.yaml --trace-data trace.csv
Chat completions API — use /v1/chat/completions instead of /v1/completions:
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--api-format chat --workload-spec workload.yaml \
--trace-header trace.yaml --trace-data trace.csv
Non-streaming with network RTT — disable SSE streaming and record network latency:
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--no-streaming --rtt-ms 2.5 --workload-spec workload.yaml \
--trace-header trace.yaml --trace-data trace.csv
Streaming and token counts
By default, observe uses streaming (SSE) and sends stream_options: {include_usage: true} to capture accurate token counts from the final SSE chunk. Non-streaming mode (--no-streaming) parses the full response body instead. Both modes extract finish_reason from server responses.
Prefix sharing
When the workload spec defines prefix groups, observe builds deterministic prefix strings from a fixed vocabulary, seeded by the RNG seed and group name. This activates the server's prefix cache for realistic KV cache hit rates.
Before dispatching requests, observe sends a single calibration request to measure the server's tokens-per-word ratio (typically 1.5–1.7 for BPE tokenizers). Prefix word counts are then scaled so the server tokenizes them to approximately the target prefix_length in the spec — matching what blis run simulates. The calibration result is logged at startup:
If calibration fails (server unreachable, timeout, or abnormal ratio), observe falls back to 1:1 word-to-token mapping with a warning.
Session support
If the workload spec contains session clients, observe runs in closed-loop mode: each completed request may trigger follow-up requests from the session manager, interleaved with pre-generated arrivals by arrival time.
System Prewarming¶
When targeting a cold system (fresh vLLM restart, new EPP connections), the first 30-60s of requests typically see 20-25x worse latency due to CUDA kernel compilation, memory allocator priming, and connection establishment. At high target rates, this cold-start transient can even trigger queue collapse.
The --prewarm-duration flag addresses this by running a fixed priming phase before measurement begins:
./blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--prewarm-duration 60s \
--workload chatbot --rate 20 --num-requests 500 \
--trace-header trace.yaml --trace-data trace.csv
The prewarm phase sends small, fixed requests (256 input tokens, 64 output tokens) at low concurrency (4 simultaneous requests). This exercises the full request path without risking overload, regardless of the real workload's rate or token sizes. Prewarm requests never appear in trace_data.csv.
Which warmup flag? Use --prewarm-duration when targeting a cold system (fresh restart, new connections). Use --warmup-requests only if you need to exclude initial requests for other reasons (e.g., rate ramp-up). You typically don't need both.
blis replay¶
Replays a captured TraceV2 file through the BLIS discrete-event simulator. Instead of generating synthetic requests, replay loads real request timing and token counts from the trace.
Replay-Specific Flags¶
| Flag | Type | Default | Description |
|---|---|---|---|
--trace-header |
string |
"" |
Path to TraceV2 header YAML (required) |
--trace-data |
string |
"" |
Path to TraceV2 data CSV (required) |
--results-path |
string |
"" |
File to write SimResult JSON for blis calibrate consumption |
--model |
string |
"" |
LLM name (required) |
--trace-output |
string |
"" |
Export replay results as TraceV2 files (<prefix>.yaml + <prefix>.csv); header mode: "replayed" |
Replay also accepts all shared simulation config flags (--latency-model, --total-kv-blocks, --max-num-seqs, etc.) — the same flags available in blis run. See Configuration for the full list.
Re-supply capacity-affecting flags identically on replay
Flags that shape KV-cache capacity are re-supplied on the replay CLI, not round-tripped through the trace header — replay recomputes capacity from --model and these flags exactly as blis run does. In particular, --kv-cache-dtype must be passed identically on replay (e.g. --kv-cache-dtype fp8) to reproduce the run's KV-block count and per-request metrics (INV-13). This mirrors --gpu-memory-utilization and the auto-computed --total-kv-blocks (the header records those only informationally). It is deliberately not the header-authoritative model used by --kv-offload-config — a KV dtype only scales a byte width, it drives no runtime mechanism.
How Replay Differs from blis run¶
| Aspect | blis run |
blis replay |
|---|---|---|
| Request source | Generated from workload spec or CLI distributions | Loaded from TraceV2 CSV |
| Arrival times | Synthesized by arrival process (Poisson, etc.) | Exact timestamps from trace |
| Token counts | Sampled from distributions | Actual observed values |
| Horizon | From --horizon flag or spec |
Auto-computed as 2x max arrival time (override with --horizon) |
| Output format | Full MetricsOutput JSON |
SimResult JSON array (request_id, ttft_us, e2e_us, input_tokens, output_tokens) |
| Session support | Session manager creates follow-ups | Session structure encoded in trace (no manager needed) |
| Trace export | --trace-output (header mode: "generated") |
--trace-output (header mode: "replayed"); see restrictions below |
--trace-output faithfully re-exports closed-loop follow-up rounds (#1630)
A closed-loop replay's --trace-output captures every round — the round-0
requests and the follow-up rounds the session manager / pool driver generate
during the run (previously only the round-0 wave was exported, #1621):
- Accumulate corpus (
session_context_growth: accumulate— e.g. ablis convert otel/blis convert wekacorpus): the re-export re-derives per-roundinput_tokensdeltas andinput_tokens_resetcompaction markers from the replayed absolute inputs, setssession_context_growth: accumulateon the header, and records the per-round think time. Re-replaying it (closed-loop) reproduces the original run's per-request per-round input and metrics (INV-13 round-trip). - Non-accumulate multi-round corpus: the re-export carries absolute per-round
input_tokens(follow-ups inherit their session'sprefix_group/prefix_length, so the prefix is not double-counted) plus the capturedthink_time_us. - Pool mode (
--concurrent-sessions N): the re-export is a complete session corpus — all sessions (originals + clones) with every round, as a closed-loop corpus. Replay it with--session-mode closed-loop(it is not re-cloned). Aggregate conservation metrics (completed/injected requests, total input/output tokens) reproduce; per-request and cache/latency aggregates are not guaranteed — pool admission timing is data-dependent, and a non-accumulate clone's cache-busting divergence is not preserved when it shares aprefix_group.
Fixed-mode (--session-mode fixed) re-export is unchanged. Known boundary: a
length-capped round (output truncated by --max-model-len) reproduces its input
length but may diverge in prefix-cache token content across the cap boundary — the
huge-ISL agentic corpora are replayed with a large --max-model-len to avoid capping.
Latency model matters
The replay command simulates token generation using the configured latency model. For accurate calibration, choose the latency model that best matches the server's behavior. See Latency Models for guidance on selecting between roofline and trained-physics modes.
Replaying OTel agentic traces¶
BLIS can replay captured OpenTelemetry agent traces (e.g. the Exgentic
agent-llm-traces dataset) as a fixed pool of N concurrent closed-loop sessions.
Each session is a linear chain of LLM calls: BLIS waits for call N to complete,
applies the recorded think/tool gap, then sends call N+1 — whose prompt reuses
call N's entire prompt plus its output plus new tokens (a strictly-growing shared
prefix for realistic KV-cache reuse).
The dataset ships as Parquet, so first explode it to per-session OTel JSON:
import pyarrow.parquet as pq, json, os
os.makedirs("otel_json", exist_ok=True)
t = pq.read_table("data/train-00001-of-00039.parquet").to_pylist()
for i, row in enumerate(t):
json.dump({"spans": row["spans"]}, open(f"otel_json/trace_{i:04d}.json", "w"))
Then convert and replay:
# Convert the trace corpus to a TraceV2 pair (accumulate = growing shared prefix).
blis convert otel --input otel_json --trace-output corpus \
--context-growth accumulate --max-think-time 15s
# Replay a fixed pool of 8 concurrent sessions, 200 total (corpus duplicated to fill).
blis replay --trace-header corpus.yaml --trace-data corpus.csv \
--model qwen/qwen3-14b --concurrent-sessions 8 --total-sessions 200
The recorded source model names are pure provenance and are dropped during
conversion — all calls are simulated under --model. (They are deliberately not
written into the trace: TraceRecord.Model is routing-significant, and a name
that differs from --model would filter out every request at routing.)
--concurrent-sessions implies closed-loop session semantics; without it,
replay behaves as a standard fixed/closed-loop replay.
Driving the corpus against a real server (blis observe)¶
The same corpus can be driven against a live inference server — the
observe-side twin of blis replay --concurrent-sessions. This records real
observed timing so blis calibrate can compare the real server against the
simulator over the identical session set:
# Drive the corpus as a fixed pool of 8 concurrent sessions, 200 total, against
# a live server, recording observed timing into observed.{yaml,csv}.
blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--corpus-header corpus.yaml --corpus-data corpus.csv \
--concurrent-sessions 8 --total-sessions 200 \
--trace-header observed.yaml --trace-data observed.csv
--corpus-header / --corpus-data name the input corpus (the pair produced
by convert otel); --trace-header / --trace-data remain the output
observed trace. Corpus-mode (--concurrent-sessions > 0) is mutually exclusive
with the spec-mode inputs (--workload, --workload-spec, --rate,
--concurrency) — a corpus is the workload. --max-concurrency is
auto-raised to the pool size if set lower, so the pool is never throttled. The
pool self-drains all --total-sessions sessions (observe bounds the run by the
session count, not a clock). Because of that self-draining, the spec-mode
bounding flags --horizon and --num-requests have no meaning in corpus-mode
and are rejected rather than silently ignored — size the run with
--total-sessions.
--shuffle-corpus works in corpus-mode here too, and draws the same seeded
permutation as blis replay --shuffle-corpus (both salt the master --seed
identically). So observe and replay of one corpus under the same --seed select
the identical subset and admission order — exactly what you want when calibrating
the simulator against the real server over a matched session set. See the
subsetting tip below for the selection semantics.
Session-id header for session-aware routing¶
Each round is an independent HTTP request, so a session-aware gateway/EPP
(session-affinity or predictive-least-loaded pinning) needs the session id on
the wire to keep a session's rounds on one instance (issue #1505). blis observe
emits it on the --session-id-header header — default x-session-id, matching
the current custom scorers — for every request carrying a session id:
blis observe --server-url http://localhost:8000 --model qwen/qwen3-14b \
--corpus-header corpus.yaml --corpus-data corpus.csv \
--concurrent-sessions 8 --total-sessions 200 \
--session-id-header x-session-id \
--trace-header observed.yaml --trace-data observed.csv
Set --session-id-header to match your deployment's session-id-producer, or to
the empty string to disable emission. It applies to any session-bearing request
— corpus-mode (--concurrent-sessions) and spec-mode closed-loop
(--concurrency) alike. Closed-loop replay is the wire producer of this header:
real captured clients carry the session in telemetry, not on the request.
To calibrate real vs simulated over the same corpus:
blis replay --trace-header corpus.yaml --trace-data corpus.csv \
--model qwen/qwen3-14b --concurrent-sessions 8 --total-sessions 200 \
--results-path sim.results.json
blis calibrate --trace-header observed.yaml --trace-data observed.csv \
--sim-results sim.results.json --report calibration.json
By default the pool is self-draining: it runs until all --total-sessions
sessions complete, regardless of how many waves that takes. For blis replay
corpus-mode you may pass --horizon to impose a hard wall-clock cap on the
simulated run — sessions still queued when the cap is reached are reported as
un-admitted (a warning is logged) rather than silently dropped. blis observe
corpus-mode has no such cap: it always self-drains on session count and rejects
--horizon / --num-requests (see above), since a live run is bounded by the
corpus, not a clock.
Subsetting a large corpus: add --shuffle-corpus
--total-sessions N on its own is deterministic, ordered selection, not a
random sample: with N below the corpus size it replays the first N
sessions in file order and never touches the tail — a biased sample for a
time-sorted or harness-grouped corpus. To draw a representative subset, add
--shuffle-corpus, which applies a seeded Fisher-Yates permutation (reproducible
from --seed, on a stream independent of token generation) before selection:
--total-sessions N --shuffle-corpus yields a seeded-random N-of-corpus
subset, while --total-sessions ≥ corpus simply randomizes the admission order
with every session still running.
Size the context window to the trace
Real agentic traces (e.g. Exgentic agent-llm-traces) often carry very large
prompts — tens of thousands to well over 100K input tokens per call, growing
further each round as the shared prefix accumulates — which can exceed a
model's default --max-model-len (e.g. qwen3-14b defaults to ~41K). If the
run completes cleanly but reports completed_requests: 0 with a matching
dropped_unservable count, the prompts didn't fit: raise --max-model-len to
cover the largest round, and scale --total-kv-blocks up proportionally so
the KV cache can hold the growing sessions.
Weka CC traces (blis convert weka)¶
The SemiAnalysis WekaTrace datasets are a second agentic-trace family. Unlike the OTel/Exgentic corpus they ship as JSONL — one proxy session per line — so no Parquet explosion step is needed:
# Convert (one session per JSONL line) to the same TraceV2 corpus format.
blis convert weka --input traces.jsonl --trace-output corpus \
--context-growth accumulate --max-think-time 0
# Replay closed-loop (or as a concurrent pool, exactly as for OTel above).
blis replay --trace-header corpus.yaml --trace-data corpus.csv \
--model qwen/qwen3-14b --session-mode closed-loop --max-model-len 1000000
The reader filters each session's requests[] to the linear main-agent stream —
type:"subagent" groups are skipped (deferred to a later PR); their wall-clock is
absorbed into the following main turn's think gap. Per-round pure client think is
recomputed as max(0, t_i − t_{i-1} − api_time_{i-1}) between consecutive main turns
(carried in the think_time_us column, --max-think-time default 0 = uncapped,
since Weka gaps are genuine away-from-keyboard times). The column is non-lossy
(#1608): a genuinely-zero recomputed think (an overlapping turn) is recorded as &0,
distinct from a not-recorded (empty) cell, so an all-overlap session replays with the
recorded zeros rather than degrading to arrival-gap think. The recorded claude-*
model names are dropped during conversion (same routing-safety reason as OTel).
Context compaction is represented (#1609)
Weka input token counts are very large (p50 ≈ 110K, p90 ≈ 395K), so the
--max-model-len / --total-kv-blocks sizing warning above applies with extra
force. Real Claude Code traffic compacts/trims context constantly — ~30% of
rounds on the full 051926 dataset have in_N < in_{N-1}+out_{N-1} (the model
summarized or trimmed). The shared agentic-trace encoder emits a per-round
input_tokens_reset marker (the recorded absolute) on exactly those non-monotone
rounds, and accumulate closed-loop replay re-seeds its growing buffer to that
absolute at the compaction boundary. The reconstructed cumulative input therefore
tracks the recorded total — previously the buffer could only grow, clamped the delta
to 0, and over-counted by ≈3–4× (+312% on that dataset). The re-seed intentionally
breaks strict prefix identity across the boundary (a summary is not a literal prefix of
the pre-compaction context), which also corrects the prefix-cache hit-rate over-estimate.
The marker is a trailing conditional CSV column: absent for monotone sessions and for
blis run, so a trace with no compaction round replays byte-identically to before
(INV-6). Traces converted by an older build (no marker column) still over-count —
re-run convert to get compaction-aware output. (The separate think-time lossy-0
sentinel was resolved in #1608 — see the non-lossy note above.)
blis calibrate¶
Compares real observed latencies (from blis observe) against simulator predictions (from blis replay) and produces a calibration report.
Flags¶
| Flag | Type | Default | Description |
|---|---|---|---|
--trace-header |
string |
"" |
Path to TraceV2 header YAML (required) |
--trace-data |
string |
"" |
Path to TraceV2 data CSV (required) |
--sim-results |
string |
"" |
Path to SimResult JSON from blis replay (required) |
--report |
string |
"" |
Path to write calibration report JSON (required) |
--warmup-requests |
int |
-1 |
Requests to exclude from comparison (-1 = use trace header value, 0 = include all) |
--network-rtt-us |
int64 |
-1 |
Network RTT in microseconds added to sim-side latencies (-1 = use trace header value) |
--network-bandwidth-mbps |
float64 |
0 |
Network bandwidth in Mbps for upload/download delay (0 = no delay) |
--itl-data |
string |
"" |
Path to ITL CSV from blis observe --record-itl to include ITL metric in the calibration report |
--throughput-tolerance-pct |
float64 |
0 |
Within-tolerance verdict (percent) on real-vs-sim output-token throughput (#1647). Verdict emitted only when > 0 |
--num-gpus |
int |
0 |
GPU count (TP×PP×DP×instances) for per-GPU throughput normalization (#1647). Per-GPU fields emitted only when > 0 |
--replay-mode |
string |
"fixed" |
Replay mode the SimResults were produced under: fixed or closed-loop (#1647). The throughput makespan is valid only for fixed; calibrate refuses the throughput block under closed-loop (or a delta corpus). Operator-affirmed since calibrate cannot observe the blis replay --session-mode choice (durable auto-stamp: #1652) |
Sentinel defaults
The --warmup-requests and --network-rtt-us flags use -1 as a sentinel meaning "read the value from the trace header." This allows the calibration to automatically use the warmup count and RTT recorded during observation. Pass 0 explicitly to override (include all requests or apply no RTT correction).
Interpreting the Calibration Report¶
The calibration report JSON contains four sections:
trace_info — Summary of the input data:
{
"num_requests": 100,
"warm_up_excluded": 5,
"matched_pairs": 95,
"token_mismatches": 2,
"itl_dropped": 3,
"duration": "2m30s"
}
matched_pairs: Requests matched by ID between trace and sim resultstoken_mismatches: Pairs where observed and simulated token counts differ (indicates potential data quality issues)itl_dropped: Requests dropped from ITL computation because all inter-chunk deltas were negative (clock skew); only present when greater than 0
metrics — Per-metric comparison. Keys are ttft, e2e, and (if --itl-data was supplied) itl. Each entry has two sub-objects and a top-level count:
{
"ttft": {
"workload_level": {
"real_mean": 1534.2,
"sim_mean": 1498.0,
"mean_error": -36.2,
"mean_percent_error": 0.024,
"real_median": 1234.5,
"sim_median": 1200.0,
"median_error": -34.5,
"median_percent_error": 0.028,
"real_p50": 1234.5,
"sim_p50": 1200.0,
"real_p90": 3200.0,
"sim_p90": 3150.0,
"real_p95": 4000.0,
"sim_p95": 3950.0,
"real_p99": 4567.8,
"sim_p99": 4500.0
},
"request_level": {
"mape": 0.05,
"pearson_r": 0.95,
"bias_direction": "under-predict",
"quality": "good"
},
"count": 95
},
"e2e": { ... }
}
The report uses two levels of analysis because they catch different problems. Workload-level aggregates (mean, median, percentiles) reveal systematic bias — the simulator consistently over- or under-predicting. Request-level prediction quality (MAPE, Pearson) captures per-request variance — how tightly the simulator tracks individual latencies regardless of any overall offset.
workload_level fields:
| Field | Meaning |
|---|---|
real_mean / sim_mean |
Arithmetic mean of observed / simulated latencies (µs) |
mean_error |
sim_mean - real_mean — positive = over-predict, negative = under-predict |
mean_percent_error |
|mean_error| / real_mean — absolute relative error on the mean |
real_median / sim_median |
Median latency aliased from P50 (µs) |
median_error |
sim_median - real_median |
median_percent_error |
|median_error| / real_median |
real_p50/p90/p95/p99 |
Real (observed) latency percentiles (µs) |
sim_p50/p90/p95/p99 |
Simulated latency percentiles (µs) |
request_level fields:
| Field | Meaning |
|---|---|
mape |
Mean Absolute Percentage Error across matched request pairs (lower is better) |
pearson_r |
Pearson correlation coefficient (closer to 1.0 is better) |
bias_direction |
over-predict, under-predict, or neutral |
quality |
Rating: excellent, good, fair, or poor |
Top-level: count — number of matched request pairs used for this metric.
config_match — Tracks which simulator config parameters matched the observed server config (currently reports matched and defaulted arrays).
known_limitations — Documents known sources of sim/real divergence (batch step granularity, synthetic prefix tokens, speculative decoding).
throughput (optional, #1647) — Aggregate throughput comparison over the request_id-matched set, further restricted to real records with status == "ok" and a positive client-frame duration (a superset filter of the latency leg: latency matches on request-ID alone, throughput additionally drops failed and zero-duration records). Present automatically whenever (a) replay-mode provenance confirms the makespan is valid — --replay-mode fixed (the default) and no delta-corpus header marker (see the validity boundary below) — and (b) it is derivable from the required inputs (a positive real and sim makespan and non-zero matched output tokens); omitted otherwise, so a refused or non-derivable run keeps the legacy report shape and no Inf/NaN value is ever written. Matched requests that failed in the real trace but completed in the sim are excluded and reported via a stderr warning unconditionally — including when every matched request failed (so the block collapses to nil) — so the completion-rate mismatch is never swallowed even though the throughput numbers cannot reflect it.
{
"matched_requests": 95,
"real_runtime_sec": 148.2,
"sim_runtime_sec": 132.7,
"real_output_tokens": 12040,
"sim_output_tokens": 12040,
"real_output_tokens_per_sec": 81.2,
"sim_output_tokens_per_sec": 90.7,
"output_tokens_per_sec_error": 9.5,
"output_tokens_per_sec_percent_error": 0.117,
"real_requests_per_sec": 0.64,
"sim_requests_per_sec": 0.72,
"requests_per_sec_error": 0.08,
"requests_per_sec_percent_error": 0.117,
"num_gpus": 4,
"real_output_tokens_per_sec_per_gpu": 20.3,
"sim_output_tokens_per_sec_per_gpu": 22.7,
"tolerance_pct": 15,
"within": true
}
- Makespan (both sides in the client frame).
real_runtime_sec= latest client last-chunk − earliest client send;sim_runtime_sec= latest (client send + client-frame sim E2E) − earliest send, where the sim E2E is normalized with the same network shift (+ network-rtt-us + upload + download) the latency comparison applies. Both share the identical earliest-send origin, so the block isolates the batching/contention difference between real and sim rather than a frame offset — and stays consistent with thee2emetric block. - Why throughput and latency can disagree. A sim can track per-request latency ordering well (high Pearson) yet mispredict aggregate throughput if it mismodels batching/contention. This block is the throughput counterpart to the latency verdicts.
num_gpus/*_per_gpuappear only with--num-gpus N(the standard normalized benchmark metric; lets a 70B/TP4 run be compared against a 7B/TP1 run). Since one GPU count divides both sides, the per-GPUpercent_errorequals the rawpercent_error— the per-GPU figures add comparability, not a new tolerance dimension.tolerance_pct/withinappear only with--throughput-tolerance-pct P;withintests the raw output-token-throughput percent error againstP. Units:*_percent_errorfields are stored as a fraction (0.117= 11.7%, the BLIS-wide convention shared withmean_percent_erroretc.), whereastolerance_pctis a percentage (15= 15%). The verdict compares them correctly (percent_error × 100 ≤ tolerance_pct); a consumer reading both must apply the ×100.- Only
status == "ok"records contribute (matching the goodput numerator). Failed/timed-out records are excluded on both the numerator and the makespan. If the sim completes only a fraction of the eligible real requests, calibrate logs a coverage warning (real throughput would otherwise be measured over the survivor subset and understate the gap); cross-checktrace_info.matched_pairsagainstnum_requests. - Validity boundary (enforced by refusal). The reconstructed sim makespan (
send + simE2E) is a physical sim timeline only for fixed-mode replay (the standard observe→replay→calibrate path, arrivals pinned from the trace). Under closed-loop / concurrent-session replay the sim regenerates the arrival schedule (round N+1 depends on the sim's completion of round N), so the real send schedule is not the sim's arrival schedule and the makespan-based throughput is meaningless.calibraterefuses the throughput block unless provenance confirms fixed mode. The block is emitted only when--replay-mode fixed(the default) and the trace header carries no delta-corpus marker (session_context_growth). Under--replay-mode closed-loop, or when the header betrays a delta/accumulate corpus, calibrate omits the throughput block and logs a loudREFUSEDwarning rather than emitting a silently-invalid verdict — turning the earlier documented blind spot into a fail-loud guard (R1). Because calibrate cannot itself observe theblis replay --session-modechoice (it is recorded in neither the trace header norSimResult), the operator affirms it via--replay-mode; if you replayed closed-loop, pass--replay-mode closed-loop(or simply do not request a throughput verdict). The durable fix — stamping the replay mode intoSimResultso calibrate can auto-detect it without the operator flag — is tracked in #1652.
Known Gap: Scheduling Delay¶
Scheduling delay cannot be calibrated
blis run and blis replay report scheduling delay — the time a request
waited in the simulator's internal queue before being selected for batch execution.
This appears in per-request RequestMetrics.SchedulingDelay and in the aggregate
scheduling_delay_p99_ms field of MetricsOutput (printed to stdout).
blis observe cannot record scheduling delay because real inference servers
do not expose per-request queue wait time through their HTTP APIs. A client sees
only TTFT and E2E; the server never reveals how long a request queued before
execution began. This gap is inherent and cannot be closed without server-side
instrumentation outside BLIS's scope.
What this means for calibration:
- Scheduling delay is already factored into E2E and TTFT. A real server's TTFT includes queue wait implicitly — it cannot be decomposed from execution time externally. Mathematically,
SchedulingDelay = (batch selection time) − ArrivalTimeandTTFT ≥ SchedulingDelayalways (per INV-5 Causality:arrival_time ≤ enqueue_time ≤ schedule_time ≤ completion_time).blis calibratecompares E2E and TTFT end-to-end; a good MAPE on those metrics confirms the scheduling model is correct, even without a direct scheduling delay comparison. - Do not expect
SchedulingDelayto appear as a separately calibratable field in the calibration report.blis calibratecompares E2E and TTFT only —scheduling_delay_msis not inSimResult(the formatblis replaywrites to--results-path). It is a simulator-internal diagnostic only. -
When scheduling delay matters: If
blis runorblis replayshows high scheduling delay (e.g., P99 > 500ms — exact thresholds are workload-dependent) and calibration MAPE is high, the root cause is likely queue buildup rather than the execution latency model. Enable flow control to add explicit queue depth gating that better models servers with admission backpressure, then re-calibrate:
Worked Example¶
This walkthrough demonstrates the full pipeline: define a workload, observe a real vLLM server, replay through the simulator, and interpret the calibration report.
Step 1: Define a workload¶
Create a workload spec (workload.yaml):
rate: 5.0
num_requests: 50
clients:
- id: "chat-user"
rate_fraction: 1.0
slo_class: "standard"
arrival:
process: poisson
input_distribution:
type: gaussian
params:
mean: 256
std_dev: 64
min: 32
max: 1024
output_distribution:
type: gaussian
params:
mean: 128
std_dev: 32
min: 16
max: 512
Step 2: Observe the real server¶
./blis observe \
--server-url http://localhost:8000 \
--model qwen/qwen3-14b \
--workload-spec workload.yaml \
--warmup-requests 5 \
--trace-header trace.yaml \
--trace-data trace.csv
This sends 50 requests to the server at ~5 req/s, excludes the first 5 from the trace (warmup), and writes the TraceV2 files.
Step 3: Replay through the simulator¶
./blis replay \
--trace-header trace.yaml \
--trace-data trace.csv \
--model qwen/qwen3-14b \
--latency-model roofline \
--results-path results.json
The simulator replays the same requests (arrival times, token counts) through the DES using the roofline latency model and writes per-request results.
Step 4: Calibrate¶
./blis calibrate \
--trace-header trace.yaml \
--trace-data trace.csv \
--sim-results results.json \
--report calibration.json
The calibration command matches requests by ID, applies warmup exclusion and RTT normalization from the trace header, and produces the report.
Step 5: Interpret results¶
Look for:
request_level.mape< 0.10 andrequest_level.quality="good"or"excellent"→ simulator tracks individual request latencies wellworkload_level.mean_percent_error< 0.05 → no systematic bias; mean prediction is within 5% of realityrequest_level.bias_direction="over-predict"→ simulator latencies are higher than reality (conservative)request_level.bias_direction="under-predict"→ simulator latencies are lower than reality (optimistic — may need latency model tuning)- High
token_mismatches→ data quality issue; check if the server truncated outputs
Low MAPE with high mean_percent_error indicates low per-request variance but a systematic offset — the simulator is consistently biased in one direction on every request. High MAPE with high mean_percent_error suggests widespread per-request inaccuracy with an additional systematic component; consider switching latency models.
If calibration quality is poor, try:
- Different latency model: Switch from
rooflinetotrained-physics(see Latency Models) - Adjust server config flags: Match
--max-num-seqsand--max-num-batched-tokensto the real server's settings - Increase sample size: Use more requests (
--num-requests) for statistical stability
GIE Headers for llm-d¶
When observing an llm-d cluster with Gateway Inference Extension (GIE), blis observe automatically sends two HTTP headers that GIE's Endpoint Picker (EPP) uses for admission control:
| Header | Workload spec field | Purpose |
|---|---|---|
x-gateway-inference-objective |
slo_class |
Name of an InferenceObjective CRD on the target cluster. EPP looks up this CRD and reads its spec.priority integer for queue ordering and shedding. |
x-gateway-inference-fairness-id |
tenant_id |
Tenant key for per-tenant fair-share scheduling. Fairness is enforced between requests of the same priority level. |
Headers are only sent when the field is non-empty, so non-GIE servers are unaffected.
How GIE resolves priority¶
GIE does not accept a priority integer directly from the client. Instead, priority is resolved server-side through a CRD lookup:
- Client sends
x-gateway-inference-objective: critical(theslo_classvalue) - EPP looks up
InferenceObjective/criticalCRD on the cluster - EPP reads
spec.priority(e.g. 100) from the CRD - The integer priority is used for strict priority queue ordering and shedding decisions
If no matching CRD exists, EPP defaults to priority 0.
Prerequisite: deploy InferenceObjective CRDs¶
For GIE headers to have any effect, matching InferenceObjective CRDs must exist on the target cluster. The CRD names must match the slo_class values in your workload spec:
# On your Kubernetes cluster:
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceObjective
metadata:
name: critical # must match slo_class in workload spec
spec:
priority: 100 # higher = more important; negative = sheddable
poolRef:
name: my-pool
---
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferenceObjective
metadata:
name: background
spec:
priority: -10 # negative priority → shed under load (HTTP 503)
poolRef:
name: my-pool
Workload spec example¶
Set slo_class and tenant_id on your clients to activate GIE headers. The slo_class value must match an InferenceObjective CRD name on the target cluster:
clients:
- id: "realtime-api"
slo_class: "critical" # → sent as x-gateway-inference-objective header
tenant_id: "team-alpha" # → sent as x-gateway-inference-fairness-id header
- id: "batch-job"
slo_class: "background" # → GIE resolves to negative priority via CRD
tenant_id: "team-beta"
Note: The GIE API version (v1alpha2) shown above may differ on your cluster. Check your installed CRD version with kubectl get crd inferenceobjectives.inference.networking.x-k8s.io.
Direct vLLM priority injection¶
When slo_class is set, blis observe also injects a priority integer into the JSON request body. This targets vLLM directly when running without GIE in front of it. vLLM uses the opposite convention from llm-d/GAIE — lower integer = more urgent (min-heap) — so the value is computed via SLOPriorityMap.InvertForVLLM(class) = maxPriority - priority(class). With the default tier priorities, this produces: critical → 0, standard → 1, batch → 5, sheddable → 6, background → 7.
Both signals are sent on every request when slo_class is set: the x-gateway-inference-objective header (llm-d) and the priority body field (vLLM). Servers ignore what they don't use — vLLM ignores unknown headers, and llm-d's vLLM FCFS backend silently ignores the body priority field — so dual delivery is safe regardless of deployment.
Tips¶
- Warmup requests: Always use
--warmup-requestsduring observation to exclude cold-start latencies (JIT compilation, KV cache initialization) from the trace. - Network RTT: If observing a remote server, measure RTT with
pingand pass--rtt-ms. The calibrate command uses this to normalize sim-side latencies. - Reproducibility: The
--seedflag controls workload generation RNG. Same seed + same spec = same request sequence. - Graceful shutdown: Press Ctrl+C during observation to stop gracefully — in-flight requests complete and all recorded data is written to the trace files.
- Large workloads: Use
--max-concurrencyto limit in-flight requests and avoid overwhelming the server.