gdput

Effective inference under real constraints.

gdput (pronounced "goodput") is an open technical reference for measuring, modeling, and optimizing AI inference effectiveness under service-level objectives and compute constraints.

Maximum raw throughput is not maximum effectiveness. The best inference system is the one that produces the most useful output while satisfying defined SLOs and using resources efficiently. gdput publishes the definitions, methodologies, measurements, and decision frameworks used to reason about that.

The research program behind the site: SLO as a function of inference configuration — treating TTFT, TPOT, and goodput as measurable, learnable functions of engine choice, runtime flags, parallelism, hardware, and workload, rather than quantities rediscovered by trial and error for every deployment.

Start here

Definitions

Canonical, measurable definitions: inference goodput, effectiveness, SLO attainment, TTFT, TPOT.

Research

The research agenda: modeling SLOs as functions of inference configuration.

Papers

Technical papers published as reproducible packages — PDF, Markdown, data, and code.

Benchmarks

SLO-aware benchmark results with raw data, configurations, and methodology.

Methods

How gdput measures: benchmark methodology, workload specifications, metric conventions.

first publications in preparation — GTP-001 coming