API connection Not configured
Not checked
Sends a one-token streaming test. Does not add to your chat.
Queued requests —Configure an endpoint to see its queue
IDEAS AT INFERENCE SPEED
A thought.
A little something real.
Describe a poster, a tiny app, or an idea worth exploring.
Watch it take shape, then try it right here.
Streamed generationThinkingOutput
Waiting for generation…
Received tok/s · rolling ~1 s window · 200 ms samples · prefill and finalization excluded
Peak streamed TPSHighest sampled rate on the curve
—tok/s
Thinking device TPS—Available after completion
Output device TPS—Available after completion
Generated tokens—Elapsed 0.0 s
First response TTFT—Ready
Measurement details
Device TPS uses server execution intervals. Streamed TPS = generated tokens received after the first content batch ÷ time from first to last content batch. Both exclude prefill and finalization; TTFT and total elapsed are separate. Streamed TPS requires server token-progress metadata; text length and event counts are never used as token counts.
Raw measurement
Request settings
Disable formatting instructions to send your prompt without an added system message. The budget hint is ignored when thinking is off. Sampling otherwise uses server defaults.