Latency preference

Fast free LLM API candidates

These candidates carry the catalog's Fast latency class. Treat that classification as a discovery signal, then measure the same prompt from your own network and account.

Current shortlist3 active model candidates

Latest provider source check: 2026-07-28.

What this shortlist means

The inclusion rules are visible and intentionally narrow.

  1. 01

    Inclusion rule

    The current model record has the Fast latency class.

  2. 02

    Inclusion rule

    The provider and model are active in the catalog.

  3. 03

    Inclusion rule

    Rate-limit text remains beside the speed classification.

Important limit: Fast is not a live latency guarantee. Queueing, model load, geography, prompt size, free-tier limits and provider incidents can all change observed speed.

Matching candidates

Live catalog results with provider evidence dates and the limits you should read before signup.

Model and providerWhy it matchesEvidence and limitsNext steps
Gemini Flash
Google Gemini
gemini-flash-latest
Fast catalog classification
Account-specific no-card tier
Official sources: 2026-07-28
Not required to start the documented free tier
Project and model specific
Independent Canary pending
Qwen3.6 27B
GroqCloud
qwen/qwen3.6-27b
Fast catalog classification
No-card free plan
Official sources: 2026-07-28
Not required for the Free plan
Model and organization specific
Independent Canary pending
GPT-OSS 120B
Cerebras Inference
gpt-oss-120b
Fast catalog classification
Payment-verified starter credits
Official sources: 2026-07-28
Verified payment method required for starter credits
Plan and model specific
Independent Canary pending