gdput

Research

research agenda · version 1.0 · 2026-07-22 · gdput Research

gdput's core research program treats the service-level objective as a function of the inference configuration — and works on learning that function.

An LLM deployment's SLO metrics — TTFT, TPOT, and ultimately goodput — are determined by its configuration: the serving engine (vLLM, SGLang, TensorRT-LLM), its runtime flags (batching capacity, KV-cache fraction, CUDA graphs, chunk sizes), parallelism degrees, hardware, and workload mix. Today that surface is mostly explored by trial and error, per engine, per cluster. The program sits in the seam between two literatures: LLM serving systems, which build the configuration knobs, and configuration-performance learning, which has a decade of machinery — transfer learning, meta-learning, structured Bayesian optimization — for modeling exactly this kind of surface but has only begun pointing it at LLM inference.

Open problems the program targets

Outputs

The agenda anchors a doctoral research program (2026–2029). It is executed in public where possible: canonical definitions, versioned measurement methodology, configuration-sweep datasets, technical papers as reproducible packages, and open-source modeling tools in the gdput GitHub organization.

answer-sized research notes in preparation