Inference effectiveness
Inference effectiveness is the degree to which an inference deployment converts its resources into useful, SLO-compliant outcomes.
- Formal definition
- Effectiveness relates achieved goodput to the resources consumed producing it. It is reported as goodput per unit of a stated resource.
- Formula
effectiveness = goodput ÷ resource resource ∈ { GPU·s, watt, dollar, accelerator instance }- Units
- SLO-compliant tokens per GPU-second, per watt, or per dollar — always with the resource named.
- Includes / excludes
- Inherits the inclusion rules of goodput; the denominator counts all resources of the deployment under test, including idle capacity reserved for it.
- Example
- Two configurations deliver
870and730 tok/sgoodput on 8×H200 and 4×H200 respectively. Effectiveness is108.75vs182.5compliant tok/s per GPU — the smaller deployment is more effective despite lower goodput. - Common misuse
- Calling raw throughput-per-dollar "effectiveness" (ignores SLO compliance); comparing across different SLO sets or workloads.
- Limitations
- Draft definition; the resource denominator must be stated for figures to be comparable.
- Related terms
- inference goodput · SLO attainment
- Preferred citation
gdput Research. "Inference effectiveness." gdput definitions, v0.1, 2026. https://gdput.com/definitions/inference-effectiveness