Loading…
Frequently Asked Questions
What is this? How do I use it?
This advisor widget suggests an LLM, engine, and configuration to service a provided workload using Modal Dedicated Endpoints. You select constraints on the inference service and a metric you'd like to optimize for, and we filter every configuration we've benchmarked down to the ones that meet your latency target and rank them by your goal.
Use the displayed command to launch a Dedicated Endpoint or head to the Modal Endpoints Dashboard
How is price per token computed?
We take the GPU-hour list price of the server and divide it by the measured total token throughput (input plus output) at the recommended target request rate. This estimate is more accurate at larger scales, where each replica will be kept closer to target rate by Modal autoscaling.
What are TTFT and TTLT?
TTFT and TTLT stand for "time to first token" and "time to last token" -- the latency felt by the client between when they send a request and when they receive the first token and the final token of the response.
You can read more about these metrics in our Almanac page on benchmarking.