The gateways for trying models: OpenRouter, Together, Groq, Fireworks
- to get started
- for all of them: an account and an API key; none requires contracts to start
- when to pick it
- You've understood you need the APIs and want to decide which counter to walk up to.
pricing
| plan | checked on | |
|---|---|---|
| OpenRouter | providers' list prices with no markup + 5.5% on credit purchases | 19/07/2026 |
| Together AI | pay per token, own list price for every hosted model | 19/07/2026 |
| Groq | pay per token; on Llama 3.3 70B, $0.59/$0.79 per million tokens in/out | 19/07/2026 |
| Fireworks AI | pay per token, with a 50% discount on batch runs | 19/07/2026 |
In the OpenRouter sheet we described the one-stop shop. But OpenRouter is a particular kind of counter: an aggregator, routing your requests to other people’s providers. Next to it stand the actual providers, hosting models on their own infrastructure. The differences matter, and the choice depends on your profile more than on a ranking.
The map
OpenRouter aggregates: one key, the widest catalogue (400+ models, including the big labs’ proprietary ones), passthrough prices. It is the counter for trying and comparing: no one else puts Claude, GPT and the Chinese open models in the same call.
Together AI hosts the open models (one of the widest catalogues among direct providers, 100+) and also lets you fine-tune a model (train it further on your data) and serve it from the same API. It is the counter for serious open-weight work.
Groq runs fast: proprietary hardware designed for inference, a narrow catalogue of selected open models, generation speeds the others don’t reach. A price anchor at this date: on Llama 3.3 70B it lists $0.59/$0.79 per million tokens in/out (source in the table, 19/07/2026). It is the counter for those with latency as a constraint: live assistants, voice, interfaces that cannot wait. Public fine-tuning: no.
Fireworks AI aims at production: function calling, structured outputs, compliance (HIPAA, zero data retention on open models without explicit opt-in), fine-tuning and serving on the same platform. It is the counter for shipping a product with requirements.
How to choose, by profile
If you are learning or comparing, OpenRouter, no hesitation: the wide catalogue is exactly its trade, and the free variants zero out the first experiments. If you have picked an open model and are building on it, Together or Fireworks, the latter favoured if you have compliance requirements or structured outputs in production. If your product lives or dies on response speed, Groq, accepting the narrow catalogue. And if you only use one lab’s models, the lab’s direct API remains the main road: one less intermediary.
On data, the operating rule is to read the chosen provider’s policy page before sending real work: retention and training policies differ and change, and they should be checked at the date, not remembered.
What this comparison lacks
Latencies measured by us: the speed figures in circulation come from the providers themselves or from third parties, and our comparative test bench (the batch evaluator in the scaffolding plan, SC-03) is not built yet. When it exists, this comparison gets updated with in-house measurements, date and conditions declared. Until then, the prices in the table carry their verification date and speed remains a manufacturer’s claim.