Which LLMs fit on an RTX 4090

As of the data snapshot of , the memory of GeForce RTX 4090 is sufficient by calculation for 101 of 1,092 catalogued configurations, 18 of them tight. Assumptions: 8,192 context tokens, concurrent requests 1, reserve 10%. 21.6 GiB remain usable. For 686 configurations a complete memory value is still missing.

The first test candidate is Devstral-Small-2-24B-Instruct-2512 in Q4_K_M with llama.cpp. This configuration needs up to 18.1 GiB, estimated at 37 to 52 tokens per second. Runtime status llama.cpp: Build b10964 from Sep 14, 2026 (release notes). Suitability: no independent data. Memory and speed are calculated, not measured. Only a test with your own questions shows whether a model suits your task.

torck LLM Finder

Your requirements

400 computable models · 1,092 configurations · 93 hardware options · Data snapshot

Straight to the result

Start with

Your hardware is known. The finder checks model configurations for this device.

Your model is known. The finder checks which listed hardware can run it under these settings.

Describe the task and requirements. The finder filters configurations using these inputs.

Select two to four configurations for a side-by-side comparison.

Select up to four configurations

Search by model, quantisation or runtime. 1,092 configurations are available.

Nothing selected yet. Search above for a model to compare configurations.

The task sets the speed guide value for the shortlist. For agents, the finder also checks support for external tools. The task does not filter models by answer quality.

Hardware undecided: start with a 48 GiB planning profile

Another 681 models are in the catalogue and are not calculated: 402 without a quantized edition at an allow-listed account, 254 no reason recorded, 20 not a language model choice, 5 format edition of another model. Open the model overview
Tokens per request, including input, history and planned output.
Number of requests processed at the same time.
Memory and serving assumptions
Missing license information is marked as unresolved.
Select a language. Without an independent test, suitability remains unverified.
The runtime executes the model. Its version, format and hardware support must align.
Quantization sets how many bits store each model value. It affects memory use and may affect quality.

Workspace is additional memory for intermediate results and the runtime. The range is an estimate. It depends on how many tokens are processed together and which runtime you use. Image processing may require additional memory.

The KV cache stores intermediate request states. Context length and concurrent requests increase its memory use.

GPU memory and CPU RAM may be used together, depending on the configuration. Actual allocation may differ.

Results 1092

20 of 1,092 results, page 1 of 55.

Calculated memory requirement: Fits the memory estimate: 83 · Fits with little headroom: 18 · At the memory limit: 5 · Memory requirement incomplete: 686 · Needs more memory: 300 · Data snapshot

What the symbols mean

  • clear: Memory is sufficient by calculation, or the figure beside it is measured.
  • with a caveat: Tight, at the memory limit, or only with CPU offload. The sentence beside it names the case.
  • not possible: Memory is missing, or runtime and hardware do not fit together.
  • open: A value is missing from the data snapshot, so it cannot be calculated.
  • estimated: Calculated, not measured.
  • measured: Measured, source and origin are stated in the details.

Models for an initial test

Guide value for this task: 10 tokens per second. That is an assumption by torck, no sourced threshold exists for it.

  • First test candidate

    Devstral-Small-2-24B-Instruct-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Nov 28, 2025 huggingface.co / Devstral-Small-2-24B-Instruct-2512

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data

    Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.

    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher

    Test suitability yourself

  • Next place in the ranking

    Magistral-Small-2509
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Sep 12, 2025 huggingface.co / Magistral-Small-2509

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data

    Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.

    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.

    The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Magistral-Small-2509 at the publisher

    Test suitability yourself

  • Next place in the ranking

    Mistral-Small-3.2-24B-Instruct-2506
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Jun 19, 2025 huggingface.co / Mistral-Small-3.2-24B-Instruct-2506

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: vendor data only

    Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.

    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task
    Vendor claim

    MMLU Pro: 69.1%

    Figure for the model variant. No measurement is available for this quantization.Vendor-reported with five examples and chain of thought. Not independently reproduced.Hugging Face Hub · Jun 20, 2025

    Test suitability yourself

The whole list follows one ranking. Configurations with verified data come first. Then it counts whether the estimated minimum speed at the selected context reaches the guide value for your task. After that the larger size class wins, within one size class the model whose repository was created later, then the higher parameter count, finally memory headroom and verified runtime data. The order says nothing about answer quality.

Decision brief and test plan

Decision brief and test plan

Data snapshot: Sep 19, 2026

Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.

This selection provides candidates for a practical test. Memory and speed are estimates. Ranking does not establish task quality or production readiness.

Start with
I have hardware
Task
General
Application language
Any
Number of identical devices
1
Distribution
Single device
KV cache: bits per value
16
KV cache: location
GPU / unified memory
Additional memory (workspace), minimum per device in GiB
1
Additional memory (workspace), maximum per device in GiB
3
Host RAM (GiB)
64
Consider CPU offload
No
Process images / scanned documents
No
Require tool use
No
For business use
No
Independent tests for the selected language
No

Calculation: memory by memory-envelope-v1, ranking by shortlist-rank-v3.

Devstral-Small-2-24B-Instruct-2512

GeForce RTX 4090 · Q4_K_M · llama.cpp b10964

Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.

  • Fits the memory estimate, 3.5 GiB stay free
  • Estimated 37 to 52 tokens per second at 8,192 tokens of context
  • Suitability: no independent data

Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB

Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.

Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.

Magistral-Small-2509

GeForce RTX 4090 · Q4_K_M · llama.cpp b10964

Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.

  • Fits the memory estimate, 3.5 GiB stay free
  • Estimated 37 to 52 tokens per second at 8,192 tokens of context
  • Suitability: no independent data

Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB

Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.

Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.

Mistral-Small-3.2-24B-Instruct-2506

GeForce RTX 4090 · Q4_K_M · llama.cpp b10964

Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.

  • Fits the memory estimate, 3.5 GiB stay free
  • Estimated 37 to 52 tokens per second at 8,192 tokens of context
  • Suitability: vendor data only

Required per device: 15.6 to 18.1 GiB · Available per device: 21.6 GiB

Estimated generation: 37 to 52 tokens/s. Measure time to first token and throughput under load separately.

Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.

Acceptance with your own tasks

For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.

Suggested pilot: 20 representative cases with expected answers defined in advance, including failure cases and questions with no supported answer. Use the same prompts and documents for every candidate. Set minimum quality and maximum waiting time before testing. Record task scores, citation errors, time to first token and total time per case. Repeat under planned concurrency.

The result link opens this calculation with its assumptions and sources. Catalogue updates may change later results. Also save the JSON export with sources and snapshot for an auditable record.

Local deployment and API: your cost estimate

All amounts are your own assumptions in EUR excluding tax. Fields remain on this device. Downloading the decision brief includes them in the file. Enter 0 for costs that do not apply. Compare only options meeting your quality, availability and data handling needs.

Local/month = purchase ÷ useful life + W ÷ 1000 × hours × electricity price + operations. API/month = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Account separately for retrieval, storage, redundancy and API base fees. This estimate checks neither capacity nor comparable answer quality.

Complete all fields with valid values to compare costs.

Applies to all results

Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.

Speed is estimated. Only a test with your own examples shows whether a model handles your tasks well. Memory values are calculated, not measured. The license information does not replace a legal review for your intended use.

Test suitability for your task

For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.

How the finder calculates

Upper bound: memory bandwidth divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache at the assumed fill level. The range is 60 to 80 percent of this upper bound. This is a rule of thumb, not a guarantee of accuracy.

Bytes per weight come from the file size of the model where the catalogue has it, otherwise from the bit width of the quantization. The memory demand uses the same figure.

The range applies to one request, without processing the prompt. Which runtime is meant stands on each card.

Prompt processing before the first answer is not estimated. Backend, driver and build can change the real rate considerably.

“Officially built” says nothing about speed or correct output.

All 1,092 configurations

  • Devstral-Small-2-24B-Instruct-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Nov 28, 2025 huggingface.co / Devstral-Small-2-24B-Instruct-2512

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Magistral-Small-2509
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Sep 12, 2025 huggingface.co / Magistral-Small-2509

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.

    The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Magistral-Small-2509 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Mistral-Small-3.2-24B-Instruct-2506
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Jun 19, 2025 huggingface.co / Mistral-Small-3.2-24B-Instruct-2506

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: vendor data only
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task
    Vendor claim

    MMLU Pro: 69.1%

    Figure for the model variant. No measurement is available for this quantization.Vendor-reported with five examples and chain of thought. Not independently reproduced.Hugging Face Hub · Jun 20, 2025

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Magistral-Small-2506
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Jun 4, 2025 huggingface.co / Magistral-Small-2506

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Magistral-Small-2506 the catalogue holds no publisher measurement for the task General.

    The catalogue names 23.6 billion parameters, 40,960 tokens of context length and the creation date Jun 4, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Magistral-Small-2506 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Mistral-Small-3.1-24B-Instruct-2503
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Mar 11, 2025 huggingface.co / Mistral-Small-3.1-24B-Instruct-2503

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 57 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Mistral-Small-3.1-24B-Instruct-2503 the catalogue holds no publisher measurement for the task General.

    The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Mar 11, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Mistral-Small-3.1-24B-Instruct-2503 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Codestral-22B-v0.1
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on May 29, 2024 huggingface.co / Codestral-22B-v0.1

    • Fits the memory estimate, 4.0 GiB stay free
    • Estimated 38 to 54 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.2 to 17.6 GiB
    Available per device
    21.6 GiB
    Response speed

    About 43 to 61 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Codestral-22B-v0.1 the catalogue holds no publisher measurement for the task General.

    The catalogue names 22.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2024. These figures bound the fit, they do not prove it.

    Open the model card of Codestral-22B-v0.1 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.

    Discuss a test

  • Ministral-3-14B-Instruct-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Instruct-2512

    • Fits the memory estimate, 9.4 GiB stay free
    • Estimated 61 to 85 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    9.9 to 12.2 GiB
    Available per device
    21.6 GiB
    Response speed

    About 70 to 98 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-14B-Instruct-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-14B-Instruct-2512
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Instruct-2512

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 56 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-14B-Instruct-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-14B-Reasoning-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Reasoning-2512

    • Fits the memory estimate, 9.4 GiB stay free
    • Estimated 61 to 85 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    9.9 to 12.2 GiB
    Available per device
    21.6 GiB
    Response speed

    About 70 to 98 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-14B-Reasoning-2512
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-14B-Reasoning-2512

    • Fits the memory estimate, 3.5 GiB stay free
    • Estimated 37 to 52 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    15.6 to 18.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 40 to 56 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-8B-Instruct-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Instruct-2512

    • Fits the memory estimate, 12.5 GiB stay free
    • Estimated 93 to 128 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    6.9 to 9.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 111 to 154 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-8B-Instruct-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-8B-Instruct-2512
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Instruct-2512

    • Fits the memory estimate, 8.8 GiB stay free
    • Estimated 57 to 80 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    10.5 to 12.8 GiB
    Available per device
    21.6 GiB
    Response speed

    About 64 to 89 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-8B-Instruct-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.

    Discuss a test

  • Ministral-3-8B-Reasoning-2512
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512

    • Fits the memory estimate, 12.5 GiB stay free
    • Estimated 93 to 128 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    6.9 to 9.1 GiB
    Available per device
    21.6 GiB
    Response speed

    About 111 to 154 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-8B-Reasoning-2512
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512

    • Fits the memory estimate, 8.8 GiB stay free
    • Estimated 57 to 80 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    10.5 to 12.8 GiB
    Available per device
    21.6 GiB
    Response speed

    About 64 to 89 tokens per second with a short context (512 tokens).

    Suitability for your task

    For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Ministral-3-8B-Reasoning-2512
    Quantization
    BF16
    Runtime
    vLLM

    Repository created on Oct 31, 2025 huggingface.co / Ministral-3-8B-Reasoning-2512

    • Fits with little headroom, 0.4 GiB stay free
    • Estimated 30 to 43 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    18.7 to 21.2 GiB
    Available per device
    21.6 GiB
    Response speed

    About 32 to 46 tokens per second with a short context (512 tokens).

    The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.With vLLM, under load or with multi-token prediction, real values can differ considerably.
    Suitability for your task

    For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.

    Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • DeepSeek-R1-0528-Qwen3-8B
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B

    • Fits the memory estimate, 12.6 GiB stay free
    • Estimated 94 to 130 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    6.8 to 9.0 GiB
    Available per device
    21.6 GiB
    Response speed

    About 115 to 159 tokens per second with a short context (512 tokens).

    Suitability for your task

    For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.

    Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • DeepSeek-R1-0528-Qwen3-8B
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B

    • Fits the memory estimate, 9.1 GiB stay free
    • Estimated 59 to 82 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    10.2 to 12.5 GiB
    Available per device
    21.6 GiB
    Response speed

    About 66 to 92 tokens per second with a short context (512 tokens).

    Suitability for your task

    For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.

    Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • DeepSeek-R1-0528-Qwen3-8B
    Quantization
    BF16
    Runtime
    vLLM

    Repository created on May 29, 2025 huggingface.co / DeepSeek-R1-0528-Qwen3-8B

    • Fits with little headroom, 1.7 GiB stay free
    • Estimated 33 to 46 tokens per second at 8,192 tokens of context
    • Suitability: no independent data
    Memory fit
    Required per device
    17.4 to 19.9 GiB
    Available per device
    21.6 GiB
    Response speed

    About 35 to 50 tokens per second with a short context (512 tokens).

    The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.With vLLM, under load or with multi-token prediction, real values can differ considerably.
    Suitability for your task

    For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.

    The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.

    Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Test with your own tasksThe list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.

    Discuss a test

  • Qwen3-8B
    Quantization
    Q4_K_M
    Runtime
    llama.cpp

    Repository created on Apr 27, 2025 huggingface.co / Qwen3-8B

    • Fits the memory estimate, 12.6 GiB stay free
    • Estimated 94 to 130 tokens per second at 8,192 tokens of context
    • Suitability: vendor data only
    Memory fit
    Required per device
    6.8 to 9.0 GiB
    Available per device
    21.6 GiB
    Response speed

    About 115 to 159 tokens per second with a short context (512 tokens).

    Suitability for your task
    Vendor claim

    IFEval strict prompt / non-thinking: 83.0%

    Figure for the model variant. No measurement is available for this quantization.Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) · May 14, 2025

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

  • Qwen3-8B
    Quantization
    Q8_0
    Runtime
    llama.cpp

    Repository created on Apr 27, 2025 huggingface.co / Qwen3-8B

    • Fits the memory estimate, 9.1 GiB stay free
    • Estimated 59 to 82 tokens per second at 8,192 tokens of context
    • Suitability: vendor data only
    Memory fit
    Required per device
    10.2 to 12.5 GiB
    Available per device
    21.6 GiB
    Response speed

    About 66 to 92 tokens per second with a short context (512 tokens).

    Suitability for your task
    Vendor claim

    IFEval strict prompt / non-thinking: 83.0%

    Figure for the model variant. No measurement is available for this quantization.Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) · May 14, 2025

    Test suitability yourself

    Why this assessment

    Show the calculation and its sources

How to check the result on an RTX 4090

Model weights, quantization and memory for active requests determine which configurations fit an RTX 4090. The calculator uses catalogue hardware data with sources in each configuration’s details.

This example uses 8,192 context tokens and one concurrent request. It includes a ten percent memory reserve and the configured runtime workspace. Increasing context or concurrency can change the selection.

Check the upper memory bound first. Then compare estimated generation speed and measure your model file on the target device. Time to first token and throughput with multiple users require separate measurements.

A memory calculation alone is insufficient for procurement. Use the decision brief to test the same business tasks with each candidate. Check licences, runtime support and outstanding evidence per configuration.

Model information

The catalog lists parameters, context length, license, and source. A catalog entry does not confirm that a model will run on your hardware.

The catalogue holds 1,081 models with technical data, 3,958 more by name only and 1,154 format editions of those models. The search finds all three.

Publisher directory

The finder reads the publishers’ model accounts in full. This overview shows per account how many models it computes with data and how many it only lists so far. The individual names are in the search above.

1,081 models with data, 3,958 more listed only, from 53 publisher accounts.

For the models that are only listed, the next reconciliation run fetches the technical data. Until then you find the name through the search and the model through the link to the account.

On top of that there are 1,154 format editions. That is the same model as GGUF, MLX, FP8, NVFP4, AWQ or GPTQ. The finder calculates them as a configuration of the original and lists their name so the search finds it.

PublisherWith dataListed onlyFormat editions
01-ai6220
ai21labs1114
Aleph-Alpha7133
allenai2579220
apple5492
arcee-ai2010159
baichuan-inc585
baidu1431
ByteDance-Seed12282
CohereLabs161412
deepseek-ai68480
google6674937
HuggingFaceTB223410
ibm-granite567560
IFM12149
inclusionAI2311837
internlm96929
JetBrains31116
kakaocorp1420
LGAI-EXAONE13928
LiquidAI1623106
llm-jp72423
meta-llama62307
microsoft8323429
MiniMaxAI1670
mistralai331716
moonshotai1600
naver-hyperclovax420
NousResearch177533
nvidia29423111
occiglot1000
openai5130
openbmb226861
openGPT-X300
OpenGVLab418123
PleIAs1275
Qwen86111239
rinna42627
Salesforce101045
sarvamai526
ServiceNow-AI711
Skywork13205
stabilityai16213
stepfun-ai14108
swiss-ai1309
tencent303338
tiiuae272659
trillionlabs8172
upstage1671
utter-project20160
xai-org200
XiaomiMiMo961
zai-org557622

Open catalogue for retrieval

The data this page calculates from is available as an open interface. Every observation names its source, that source's license and the date of capture. Responses are JSON and paginated, no account is needed.

Only facts whose source permits sharing with attribution are served. The terms of the source also apply to any reuse; the exclusion rule is stated in the header. Data snapshot .

Sources and licenses

The finder takes single data points from these sources and turns them into memory and suitability assessments. The sources have not reviewed this evaluation.

  • AMD
    No data license stated · Individual facts with source link · Retrieved Sep 20, 2026 · Source documents: 15
  • Apple
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 21
  • heise online, Jan-Keno Janssen
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1
  • Hugging Face, Modellliste 01-ai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste ai21labs
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste Aleph-Alpha
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste allenai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste apple
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste arcee-ai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste baichuan-inc
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste baidu
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste ByteDance-Seed
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste CohereLabs
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste deepseek-ai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste google
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste HuggingFaceTB
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste ibm-granite
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste IFM
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste inclusionAI
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste internlm
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste JetBrains
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste kakaocorp
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste LGAI-EXAONE
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste LiquidAI
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste llm-jp
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste meta-llama
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste microsoft
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste MiniMaxAI
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste mistralai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste moonshotai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste naver-hyperclovax
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste NousResearch
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste nvidia
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste occiglot
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste openai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste openbmb
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste openGPT-X
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste OpenGVLab
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste PleIAs
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste Qwen
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste rinna
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste Salesforce
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste sarvamai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste ServiceNow-AI
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste Skywork
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste stabilityai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste stepfun-ai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste swiss-ai
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste tencent
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste tiiuae
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste trillionlabs
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste upstage
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste utter-project
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste xai-org
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste XiaomiMiMo
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face, Modellliste zai-org
    License: Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • Hugging Face Hub
    License: Apache 2.0; Gemma Terms of Use; Llama 3 Community License; Llama 3.1 Community License; Llama 3.2 Community License; Llama 3.3 Community License; MIT; MIT with the terms of the Llama base model; Publisher's own license (text at the source) · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 3970
  • Intel
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 7
  • Knoop, Holtmann (arXiv:2601.09527)
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1
  • llama-roofline, Manu Nicholas Jacob
    License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1
  • llama.cpp, The ggml authors
    License: MIT · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 10
  • Max Vyaznikov, GPU Ark (Zenodo, doi:10.5281/zenodo.20390790)
    License: CC BY 4.0 · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1
  • Meta Llama
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 2
  • MLX, MLX Contributors
    License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1
  • mlx-lm, Apple Inc.
    License: MIT · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 1
  • mlx-lm, MLX Contributors
    License: MIT · Used under license and terms of use · Retrieved Sep 17, 2026 · Source documents: 1
  • NVIDIA
    No data license stated · Individual facts with source link · Retrieved Sep 20, 2026 · Source documents: 27
  • Qwen Team, Qwen3 Technical Report (arXiv:2505.09388)
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1
  • Qwen Team
    No data license stated · Individual facts with source link · Retrieved Sep 17, 2026 · Source documents: 1
  • vLLM project
    License: Apache 2.0 · Used under license and terms of use · Retrieved Sep 18, 2026 · Source documents: 10
Calculation method

The memory calculation follows method memory-envelope-v1: weights + 2*layers*kv_heads*head_dim*kv_bytes*context*concurrency + workspace

The KV cache gets a 5% allowance for memory management, the file size of the weights 3%. The runtime workspace is set at 1.0 to 3.0 GiB. A 10% reserve of device memory stays free.

The assumptions are based on the vLLM documentation on memory optimization and are a planning assumption by torck, recorded on .

Worked example for Devstral-Small-2-24B-Instruct-2512 in Q4_K_M on GeForce RTX 4090 with 8,192 context tokens and concurrent requests 1. Weights 13.3 to 13.8 GiB, KV cache 1.3 GiB, runtime workspace 1.0 to 3.0 GiB. Together that is 15.6 to 18.1 GiB. 21.6 GiB are usable.

Computable means that the model variant has at least one configuration with a documented file size. This applies to 400 of 1,081 model variants in the catalogue. The catalogue lists the others with technical data but without a computable configuration yet.

Memory requirements are calculated, not measured. They consist of the model weights based on file size or quantization bit width, the KV cache for context and concurrent requests, and the runtime workspace. The selected reserve remains free.

Response speed is estimated, not measured. The upper bound is the memory bandwidth of the device divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache. The finder shows a range of 60 to 80 percent of this upper bound. This corridor is a rule of thumb, not a guarantee of accuracy. A number appears only for CUDA and Metal, for a single request and for models that fit entirely into device memory.

For runtime and device, each configuration names an evidence level, such as “Officially built” or “Target architecture in the official build”. An evidence level says nothing about speed or error-free output. If evidence is missing, the reason is listed under open points.

The estimate can be wrong in several cases. For MoE models, the calculation can be considerably too high. Long context, backend, driver and build, several concurrent requests, vLLM under load and offload change the real rate. For ROCm and SYCL, the finder shows no number because no verified reference measurement is available. Prompt processing before the first answer is not included.