Definitions
Canonical, measurable definitions. Each states exactly what it includes, excludes, and how to compute it — so results using them are comparable.
Inference goodput
Useful output delivered per unit time while satisfying all defined service-level objectives.
Inference effectiveness
The degree to which a deployment converts its resources into useful, SLO-compliant outcomes.
SLO attainment
The proportion of requests satisfying every defined service-level objective.
Time to first token (TTFT)
Elapsed time from request arrival to the first output token, governed by the prefill phase.
Time per output token (TPOT)
Average time between successive output tokens after the first, governed by the decode phase.
more definitions in preparation: inference configuration · config→SLO surface · effective concurrency · cost per compliant token