Which LLMs fit on an RTX 4090
As of the data snapshot of , the memory of GeForce RTX 4090 is sufficient by calculation for 101 of 1,092 catalogued configurations, 18 of them tight. Assumptions: 8,192 context tokens, concurrent requests 1, reserve 10%. 21.6 GiB remain usable. For 686 configurations a complete memory value is still missing.
The first test candidate is Devstral-Small-2-24B-Instruct-2512 in Q4_K_M with llama.cpp. This configuration needs up to 18.1 GiB, estimated at 37 to 52 tokens per second. Runtime status llama.cpp: Build b10964 from Sep 14, 2026 (release notes). Suitability: no independent data. Memory and speed are calculated, not measured. Only a test with your own questions shows whether a model suits your task.
Your requirements
400 computable models · 1,092 configurations · 93 hardware options · Data snapshotResults 1092
20 of 1,092 results, page 1 of 55.
Calculated memory requirement: Fits the memory estimate: 83 · Fits with little headroom: 18 · At the memory limit: 5 · Memory requirement incomplete: 686 · Needs more memory: 300 · Data snapshot
What the symbols mean
- clear: Memory is sufficient by calculation, or the figure beside it is measured.
- with a caveat: Tight, at the memory limit, or only with CPU offload. The sentence beside it names the case.
- not possible: Memory is missing, or runtime and hardware do not fit together.
- open: A value is missing from the data snapshot, so it cannot be calculated.
- estimated: Calculated, not measured.
- measured: Measured, source and origin are stated in the details.
Models for an initial test
Guide value for this task: 10 tokens per second. That is an assumption by torck, no sourced threshold exists for it.
First test candidate
Devstral-Small-2-24B-Instruct-2512
Repository created on Nov 28, 2025 huggingface.co / Devstral-Small-2-24B-Instruct-2512
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.
Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher
Next place in the ranking
Magistral-Small-2509
Repository created on Sep 12, 2025 huggingface.co / Magistral-Small-2509
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.
The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.
Open the model card of Magistral-Small-2509 at the publisher
Next place in the ranking
Mistral-Small-3.2-24B-Instruct-2506
Repository created on Jun 19, 2025 huggingface.co / Mistral-Small-3.2-24B-Instruct-2506
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: vendor data only
Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
Vendor claimMMLU Pro: 69.1%
Figure for the model variant. No measurement is available for this quantization.Vendor-reported with five examples and chain of thought. Not independently reproduced.Hugging Face Hub · Jun 20, 2025
The whole list follows one ranking. Configurations with verified data come first. Then it counts whether the estimated minimum speed at the selected context reaches the guide value for your task. After that the larger size class wins, within one size class the model whose repository was created later, then the higher parameter count, finally memory headroom and verified runtime data. The order says nothing about answer quality.
Decision brief and test plan
Decision brief and test plan
Data snapshot: Sep 19, 2026
Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.
This selection provides candidates for a practical test. Memory and speed are estimates. Ranking does not establish task quality or production readiness.
- Start with
- I have hardware
- Task
- General
- Application language
- Any
- Number of identical devices
- 1
- Distribution
- Single device
- KV cache: bits per value
- 16
- KV cache: location
- GPU / unified memory
- Additional memory (workspace), minimum per device in GiB
- 1
- Additional memory (workspace), maximum per device in GiB
- 3
- Host RAM (GiB)
- 64
- Consider CPU offload
- No
- Process images / scanned documents
- No
- Require tool use
- No
- For business use
- No
- Independent tests for the selected language
- No
Calculation: memory by memory-envelope-v1, ranking by shortlist-rank-v3.
Devstral-Small-2-24B-Instruct-2512
GeForce RTX 4090 · Q4_K_M · llama.cpp b10964
Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB
Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.
Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.
Magistral-Small-2509
GeForce RTX 4090 · Q4_K_M · llama.cpp b10964
Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB
Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.
Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.
Mistral-Small-3.2-24B-Instruct-2506
GeForce RTX 4090 · Q4_K_M · llama.cpp b10964
Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: vendor data only
Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB
Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.
Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.
Acceptance with your own tasks
For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.
Suggested pilot: 20 representative cases with expected answers defined in advance, including failure cases and questions with no supported answer. Use the same prompts and documents for every candidate. Set minimum quality and maximum waiting time before testing. Record task scores, citation errors, time to first token and total time per case. Repeat under planned concurrency.
The result link opens this calculation with its assumptions and sources. Catalogue updates may change later results. Also save the JSON export with sources and snapshot for an auditable record.
Local deployment and API: your cost estimate
All amounts are your own assumptions in EUR excluding tax. Fields remain on this device. Downloading the decision brief includes them in the file. Enter 0 for costs that do not apply. Compare only options meeting your quality, availability and data handling needs.
Local/month = purchase ÷ useful life + W ÷ 1000 × hours × electricity price + operations. API/month = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Account separately for retrieval, storage, redundancy and API base fees. This estimate checks neither capacity nor comparable answer quality.
Applies to all results
Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.
Speed is estimated. Only a test with your own examples shows whether a model handles your tasks well. Memory values are calculated, not measured. The license information does not replace a legal review for your intended use.
Test suitability for your task
For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.
How the finder calculates
Upper bound: memory bandwidth divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache at the assumed fill level. The range is 60 to 80 percent of this upper bound. This is a rule of thumb, not a guarantee of accuracy.
Bytes per weight come from the file size of the model where the catalogue has it, otherwise from the bit width of the quantization. The memory demand uses the same figure.
The range applies to one request, without processing the prompt. Which runtime is meant stands on each card.
- Memory bandwidth 1,010 GB/s: Max Vyaznikov, GPU Ark (Zenodo, doi:10.5281/zenodo.20390790) · Sep 17, 2026
- Memory bandwidth 1,008 GB/s: NVIDIA · Sep 17, 2026
Prompt processing before the first answer is not estimated. Backend, driver and build can change the real rate considerably.
“Officially built” says nothing about speed or correct output.
All 1,092 configurations
Devstral-Small-2-24B-Instruct-2512
Repository created on Nov 28, 2025 huggingface.co / Devstral-Small-2-24B-Instruct-2512
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.
Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher
Why this assessment
Magistral-Small-2509
Repository created on Sep 12, 2025 huggingface.co / Magistral-Small-2509
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.
The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.
Open the model card of Magistral-Small-2509 at the publisher
Why this assessment
Mistral-Small-3.2-24B-Instruct-2506
Repository created on Jun 19, 2025 huggingface.co / Mistral-Small-3.2-24B-Instruct-2506
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: vendor data only
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
Vendor claimMMLU Pro: 69.1%
Figure for the model variant. No measurement is available for this quantization.Vendor-reported with five examples and chain of thought. Not independently reproduced.Hugging Face Hub · Jun 20, 2025Why this assessment
Magistral-Small-2506
Repository created on Jun 4, 2025 huggingface.co / Magistral-Small-2506
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Magistral-Small-2506 the catalogue holds no publisher measurement for the task General.
The catalogue names 23.6 billion parameters, 40,960 tokens of context length and the creation date Jun 4, 2025. These figures bound the fit, they do not prove it.
Open the model card of Magistral-Small-2506 at the publisher
Why this assessment
Mistral-Small-3.1-24B-Instruct-2503
Repository created on Mar 11, 2025 huggingface.co / Mistral-Small-3.1-24B-Instruct-2503
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 57 tokens per second with a short context (512 tokens).
Suitability for your task
For Mistral-Small-3.1-24B-Instruct-2503 the catalogue holds no publisher measurement for the task General.
The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Mar 11, 2025. These figures bound the fit, they do not prove it.
Open the model card of Mistral-Small-3.1-24B-Instruct-2503 at the publisher
Why this assessment
Codestral-22B-v0.1
Repository created on May 29, 2024 huggingface.co / Codestral-22B-v0.1
- Fits the memory estimate, 4.0 GiB stay free
- Estimated 38 to 54 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.2 to 17.6 GiB
- Available per device
- 21.6 GiB
Response speed
About 43 to 61 tokens per second with a short context (512 tokens).
Suitability for your task
For Codestral-22B-v0.1 the catalogue holds no publisher measurement for the task General.
The catalogue names 22.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2024. These figures bound the fit, they do not prove it.
Why this assessment
Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.
Ministral-3-14B-Instruct-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Instruct-2512
- Fits the memory estimate, 9.4 GiB stay free
- Estimated 61 to 85 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 9.9 to 12.2 GiB
- Available per device
- 21.6 GiB
Response speed
About 70 to 98 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-14B-Instruct-2512 at the publisher
Why this assessment
Ministral-3-14B-Instruct-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Instruct-2512
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 56 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-14B-Instruct-2512 at the publisher
Why this assessment
Ministral-3-14B-Reasoning-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Reasoning-2512
- Fits the memory estimate, 9.4 GiB stay free
- Estimated 61 to 85 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 9.9 to 12.2 GiB
- Available per device
- 21.6 GiB
Response speed
About 70 to 98 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher
Why this assessment
Ministral-3-14B-Reasoning-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Reasoning-2512
- Fits the memory estimate, 3.5 GiB stay free
- Estimated 37 to 52 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 15.6 to 18.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 40 to 56 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher
Why this assessment
Ministral-3-8B-Instruct-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Instruct-2512
- Fits the memory estimate, 12.5 GiB stay free
- Estimated 93 to 128 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 6.9 to 9.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 111 to 154 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-8B-Instruct-2512 at the publisher
Why this assessment
Ministral-3-8B-Instruct-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Instruct-2512
- Fits the memory estimate, 8.8 GiB stay free
- Estimated 57 to 80 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 10.5 to 12.8 GiB
- Available per device
- 21.6 GiB
Response speed
About 64 to 89 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-8B-Instruct-2512 at the publisher
Why this assessment
Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.
Ministral-3-8B-Reasoning-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512
- Fits the memory estimate, 12.5 GiB stay free
- Estimated 93 to 128 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 6.9 to 9.1 GiB
- Available per device
- 21.6 GiB
Response speed
About 111 to 154 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher
Why this assessment
Ministral-3-8B-Reasoning-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512
- Fits the memory estimate, 8.8 GiB stay free
- Estimated 57 to 80 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 10.5 to 12.8 GiB
- Available per device
- 21.6 GiB
Response speed
About 64 to 89 tokens per second with a short context (512 tokens).
Suitability for your task
For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher
Why this assessment
Ministral-3-8B-Reasoning-2512
Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512
- Fits with little headroom, 0.4 GiB stay free
- Estimated 30 to 43 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 18.7 to 21.2 GiB
- Available per device
- 21.6 GiB
Response speed
About 32 to 46 tokens per second with a short context (512 tokens).
The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.With vLLM, under load or with multi-token prediction, real values can differ considerably.Suitability for your task
For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.
The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.
Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher
Why this assessment
DeepSeek-R1-0528-Qwen3-8B
Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B
- Fits the memory estimate, 12.6 GiB stay free
- Estimated 94 to 130 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 6.8 to 9.0 GiB
- Available per device
- 21.6 GiB
Response speed
About 115 to 159 tokens per second with a short context (512 tokens).
Suitability for your task
For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.
The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.
Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher
Why this assessment
DeepSeek-R1-0528-Qwen3-8B
Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B
- Fits the memory estimate, 9.1 GiB stay free
- Estimated 59 to 82 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 10.2 to 12.5 GiB
- Available per device
- 21.6 GiB
Response speed
About 66 to 92 tokens per second with a short context (512 tokens).
Suitability for your task
For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.
The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.
Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher
Why this assessment
DeepSeek-R1-0528-Qwen3-8B
Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B
- Fits with little headroom, 1.7 GiB stay free
- Estimated 33 to 46 tokens per second at 8,192 tokens of context
- Suitability: no independent data
Memory fit
- Required per device
- 17.4 to 19.9 GiB
- Available per device
- 21.6 GiB
Response speed
About 35 to 50 tokens per second with a short context (512 tokens).
The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.With vLLM, under load or with multi-token prediction, real values can differ considerably.Suitability for your task
For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.
The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.
Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher
Why this assessment
Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.
Qwen3-8B
Repository created on Apr 27, 2025 huggingface.co / Qwen3-8B
- Fits the memory estimate, 12.6 GiB stay free
- Estimated 94 to 130 tokens per second at 8,192 tokens of context
- Suitability: vendor data only
Memory fit
- Required per device
- 6.8 to 9.0 GiB
- Available per device
- 21.6 GiB
Response speed
About 115 to 159 tokens per second with a short context (512 tokens).
Suitability for your task
Vendor claimIFEval strict prompt / non-thinking: 83.0%
Figure for the model variant. No measurement is available for this quantization.Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) · May 14, 2025Why this assessment
Qwen3-8B
Repository created on Apr 27, 2025 huggingface.co / Qwen3-8B
- Fits the memory estimate, 9.1 GiB stay free
- Estimated 59 to 82 tokens per second at 8,192 tokens of context
- Suitability: vendor data only
Memory fit
- Required per device
- 10.2 to 12.5 GiB
- Available per device
- 21.6 GiB
Response speed
About 66 to 92 tokens per second with a short context (512 tokens).
Suitability for your task
Vendor claimIFEval strict prompt / non-thinking: 83.0%
Figure for the model variant. No measurement is available for this quantization.Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) · May 14, 2025Why this assessment
20 of 1,092 shown
How to check the result on an RTX 4090
Model weights, quantization and memory for active requests determine which configurations fit an RTX 4090. The calculator uses catalogue hardware data with sources in each configuration’s details.
This example uses 8,192 context tokens and one concurrent request. It includes a ten percent memory reserve and the configured runtime workspace. Increasing context or concurrency can change the selection.
Check the upper memory bound first. Then compare estimated generation speed and measure your model file on the target device. Time to first token and throughput with multiple users require separate measurements.
A memory calculation alone is insufficient for procurement. Use the decision brief to test the same business tasks with each candidate. Check licences, runtime support and outstanding evidence per configuration.
Model information
The catalog lists parameters, context length, license, and source. A catalog entry does not confirm that a model will run on your hardware.
The catalogue holds 1,081 models with technical data, 3,958 more by name only and 1,154 format editions of those models. The search finds all three.
Publisher directory
The finder reads the publishers’ model accounts in full. This overview shows per account how many models it computes with data and how many it only lists so far. The individual names are in the search above.
1,081 models with data, 3,958 more listed only, from 53 publisher accounts.
For the models that are only listed, the next reconciliation run fetches the technical data. Until then you find the name through the search and the model through the link to the account.
On top of that there are 1,154 format editions. That is the same model as GGUF, MLX, FP8, NVFP4, AWQ or GPTQ. The finder calculates them as a configuration of the original and lists their name so the search finds it.
| Publisher | With data | Listed only | Format editions |
|---|---|---|---|
| 01-ai | 6 | 22 | 0 |
| ai21labs | 11 | 1 | 4 |
| Aleph-Alpha | 7 | 13 | 3 |
| allenai | 25 | 792 | 20 |
| apple | 5 | 49 | 2 |
| arcee-ai | 20 | 101 | 59 |
| baichuan-inc | 5 | 8 | 5 |
| baidu | 14 | 3 | 1 |
| ByteDance-Seed | 12 | 28 | 2 |
| CohereLabs | 16 | 14 | 12 |
| deepseek-ai | 68 | 48 | 0 |
| 66 | 749 | 37 | |
| HuggingFaceTB | 22 | 34 | 10 |
| ibm-granite | 56 | 75 | 60 |
| IFM | 12 | 14 | 9 |
| inclusionAI | 23 | 118 | 37 |
| internlm | 9 | 69 | 29 |
| JetBrains | 3 | 11 | 16 |
| kakaocorp | 14 | 2 | 0 |
| LGAI-EXAONE | 13 | 9 | 28 |
| LiquidAI | 16 | 23 | 106 |
| llm-jp | 7 | 242 | 3 |
| meta-llama | 62 | 30 | 7 |
| microsoft | 83 | 234 | 29 |
| MiniMaxAI | 16 | 7 | 0 |
| mistralai | 33 | 17 | 16 |
| moonshotai | 16 | 0 | 0 |
| naver-hyperclovax | 4 | 2 | 0 |
| NousResearch | 17 | 75 | 33 |
| nvidia | 29 | 423 | 111 |
| occiglot | 10 | 0 | 0 |
| openai | 5 | 13 | 0 |
| openbmb | 22 | 68 | 61 |
| openGPT-X | 3 | 0 | 0 |
| OpenGVLab | 4 | 181 | 23 |
| PleIAs | 12 | 7 | 5 |
| Qwen | 86 | 111 | 239 |
| rinna | 4 | 26 | 27 |
| Salesforce | 10 | 104 | 5 |
| sarvamai | 5 | 2 | 6 |
| ServiceNow-AI | 7 | 1 | 1 |
| Skywork | 13 | 20 | 5 |
| stabilityai | 16 | 21 | 3 |
| stepfun-ai | 14 | 10 | 8 |
| swiss-ai | 13 | 0 | 9 |
| tencent | 30 | 33 | 38 |
| tiiuae | 27 | 26 | 59 |
| trillionlabs | 8 | 17 | 2 |
| upstage | 16 | 7 | 1 |
| utter-project | 20 | 16 | 0 |
| xai-org | 2 | 0 | 0 |
| XiaomiMiMo | 9 | 6 | 1 |
| zai-org | 55 | 76 | 22 |
Open catalogue for retrieval
The data this page calculates from is available as an open interface. Every observation names its source, that source's license and the date of capture. Responses are JSON and paginated, no account is needed.
- Header with snapshot, counts per collection and terms of use/wp-json/torck-llm/v1/catalog
- Entries of one collection with their current values/wp-json/torck-llm/v1/catalog/entities?kind=ModelVariant
- Single values with origin, source, date and unit/wp-json/torck-llm/v1/catalog/values
- Sources with license, status and terms of use/wp-json/torck-llm/v1/catalog/sources
- Changes since an earlier snapshot/wp-json/torck-llm/v1/catalog/changes?since=0
Only facts whose source permits sharing with attribution are served. The terms of the source also apply to any reuse; the exclusion rule is stated in the header. Data snapshot .
Sources and licenses
The finder takes single data points from these sources and turns them into memory and suitability assessments. The sources have not reviewed this evaluation.
- AMD
No data license stated · Individual facts with source link · Retrieved Sep 20, 2026 · Source documents: 15 - Apple
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 21 - heise online, Jan-Keno Janssen
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1 - Hugging Face, Modellliste 01-ai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste ai21labs
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste Aleph-Alpha
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste allenai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste apple
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste arcee-ai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste baichuan-inc
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste baidu
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste ByteDance-Seed
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste CohereLabs
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste deepseek-ai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste google
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste HuggingFaceTB
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste ibm-granite
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste IFM
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste inclusionAI
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste internlm
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste JetBrains
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste kakaocorp
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste LGAI-EXAONE
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste LiquidAI
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste llm-jp
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste meta-llama
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste microsoft
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste MiniMaxAI
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste mistralai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste moonshotai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste naver-hyperclovax
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste NousResearch
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste nvidia
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste occiglot
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste openai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste openbmb
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste openGPT-X
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste OpenGVLab
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste PleIAs
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste Qwen
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste rinna
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste Salesforce
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste sarvamai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste ServiceNow-AI
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste Skywork
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste stabilityai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste stepfun-ai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste swiss-ai
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste tencent
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste tiiuae
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste trillionlabs
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste upstage
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste utter-project
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste xai-org
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste XiaomiMiMo
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face, Modellliste zai-org
License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - Hugging Face Hub
License: Apache 2.0; Gemma Terms of Use; Llama 3 Community License; Llama 3.1 Community License; Llama 3.2 Community License; Llama 3.3 Community License; MIT; MIT with the terms of the Llama base model; Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 3970 - Intel
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 7 - Knoop, Holtmann (arXiv:2601.09527)
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1 - llama-roofline, Manu Nicholas Jacob
License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1 - llama.cpp, The ggml authors
License: MIT · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 10 - Max Vyaznikov, GPU Ark (Zenodo, doi:10.5281/zenodo.20390790)
License: CC BY 4.0 · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1 - Meta Llama
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 2 - MLX, MLX Contributors
License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1 - mlx-lm, Apple Inc.
License: MIT · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1 - mlx-lm, MLX Contributors
License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1 - NVIDIA
No data license stated · Individual facts with source link · Retrieved Sep 20, 2026 · Source documents: 27 - Qwen Team, Qwen3 Technical Report (arXiv:2505.09388)
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1 - Qwen Team
No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1 - vLLM project
License: Apache 2.0 · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 10
Calculation method
The memory calculation follows method memory-envelope-v1: weights + 2*layers*kv_heads*head_dim*kv_bytes*context*concurrency + workspace
The KV cache gets a 5% allowance for memory management, the file size of the weights 3%. The runtime workspace is set at 1.0 to 3.0 GiB. A 10% reserve of device memory stays free.
The assumptions are based on the vLLM documentation on memory optimization and are a planning assumption by torck, recorded on .
Worked example for Devstral-Small-2-24B-Instruct-2512 in Q4_K_M on GeForce RTX 4090 with 8,192 context tokens and concurrent requests 1. Weights 13.3 to 13.8 GiB, KV cache 1.3 GiB, runtime workspace 1.0 to 3.0 GiB. Together that is 15.6 to 18.1 GiB. 21.6 GiB are usable.
Computable means that the model variant has at least one configuration with a documented file size. This applies to 400 of 1,081 model variants in the catalogue. The catalogue lists the others with technical data but without a computable configuration yet.
Memory requirements are calculated, not measured. They consist of the model weights based on file size or quantization bit width, the KV cache for context and concurrent requests, and the runtime workspace. The selected reserve remains free.
Response speed is estimated, not measured. The upper bound is the memory bandwidth of the device divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache. The finder shows a range of 60 to 80 percent of this upper bound. This corridor is a rule of thumb, not a guarantee of accuracy. A number appears only for CUDA and Metal, for a single request and for models that fit entirely into device memory.
For runtime and device, each configuration names an evidence level, such as “Officially built” or “Target architecture in the official build”. An evidence level says nothing about speed or error-free output. If evidence is missing, the reason is listed under open points.
The estimate can be wrong in several cases. For MoE models, the calculation can be considerably too high. Long context, backend, driver and build, several concurrent requests, vLLM under load and offload change the real rate. For ROCm and SYCL, the finder shows no number because no verified reference measurement is available. Prompt processing before the first answer is not included.