gdput

Inference effectiveness

definition · version 0.1 (draft) · 2026-07-21 · gdput Research

Inference effectiveness is the degree to which an inference deployment converts its resources into useful, SLO-compliant outcomes.

Formal definition
Effectiveness relates achieved goodput to the resources consumed producing it. It is reported as goodput per unit of a stated resource.
Formula
effectiveness = goodput ÷ resource
resource ∈ { GPU·s, watt, dollar, accelerator instance }
Units
SLO-compliant tokens per GPU-second, per watt, or per dollar — always with the resource named.
Includes / excludes
Inherits the inclusion rules of goodput; the denominator counts all resources of the deployment under test, including idle capacity reserved for it.
Example
Two configurations deliver 870 and 730 tok/s goodput on 8×H200 and 4×H200 respectively. Effectiveness is 108.75 vs 182.5 compliant tok/s per GPU — the smaller deployment is more effective despite lower goodput.
Common misuse
Calling raw throughput-per-dollar "effectiveness" (ignores SLO compliance); comparing across different SLO sets or workloads.
Limitations
Draft definition; the resource denominator must be stated for figures to be comparable.
Related terms
inference goodput · SLO attainment
Preferred citation
gdput Research. "Inference effectiveness." gdput definitions, v0.1, 2026.
https://gdput.com/definitions/inference-effectiveness