{"id":6012,"date":"2026-09-22T01:41:24","date_gmt":"2026-09-22T01:41:24","guid":{"rendered":"https:\/\/www.torck.io\/?page_id=6012"},"modified":"2026-09-22T01:41:24","modified_gmt":"2026-09-22T01:41:24","slug":"rtx-4090","status":"publish","type":"page","link":"https:\/\/www.torck.io\/en\/llm-finder\/rtx-4090\/","title":{"rendered":"Plan LLM deployment on RTX 4090"},"content":{"rendered":"<section class=\"tk-llm notranslate\" data-no-translation data-mode=\"hardware\" data-endpoint=\"https:\/\/www.torck.io\/en\/wp-json\/torck-llm\/v1\/resolve\" data-language=\"en\" lang=\"en\" translate=\"no\">\n <section class=\"tk-llm-answer notranslate\" data-no-translation translate=\"no\" lang=\"en\" aria-labelledby=\"tk-llm-answer-title\"><h2 id=\"tk-llm-answer-title\">Which LLMs fit on an RTX 4090<\/h2><p>As of the data snapshot of <time datetime=\"2026-09-19\">Sep 19, 2026<\/time>, the memory of GeForce RTX 4090 is sufficient by calculation for 101 of 1,092 catalogued configurations, 18 of them tight. Assumptions: 8,192 context tokens, concurrent requests 1, reserve 10%. 21.6 GiB remain usable. For 686 configurations a complete memory value is still missing.<\/p><p>The first test candidate is Devstral-Small-2-24B-Instruct-2512 in Q4_K_M with llama.cpp. This configuration needs up to 18.1 GiB, estimated at 37 to 52 tokens per second. Runtime status llama.cpp: Build b10964 from Sep 14, 2026 (<a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/b10964\" target=\"_blank\" rel=\"noopener noreferrer\">release notes<\/a>). Suitability: no independent data. Memory and speed are calculated, not measured. Only a test with your own questions shows whether a model suits your task.<\/p><p class=\"tk-llm-small tk-llm-answer-links\"><a href=\"#tk-llm-results-title\">Go to the full result<\/a> \u00b7 <a href=\"#tk-llm-calculation\">Calculation method<\/a><\/p><\/section> <header class=\"tk-llm-head\"><span class=\"tk-llm-kicker\">torck LLM Finder<\/span><h2>Your requirements<\/h2><span class=\"tk-llm-stamp\">400 computable models \u00b7 1,092 configurations \u00b7 93 hardware options \u00b7 Data snapshot <time datetime=\"2026-09-19\">Sep 19, 2026<\/time><\/span><\/header>\n <p class=\"tk-llm-skip\"><a href=\"#tk-llm-results-title\">Straight to the result<\/a><\/p>  <form class=\"tk-llm-form\" method=\"get\" action=\"https:\/\/www.torck.io\/en\/llm-finder\/rtx-4090\/\">\n <fieldset class=\"tk-llm-modes\"><legend>Start with<\/legend><label><input type=\"radio\" name=\"llm[mode]\" value=\"hardware\" checked aria-describedby=\"tk-llm-help-mode-hardware\"><span>I have hardware<\/span><\/label><label><input type=\"radio\" name=\"llm[mode]\" value=\"model\"  aria-describedby=\"tk-llm-help-mode-model\"><span>I have a model<\/span><\/label><label><input type=\"radio\" name=\"llm[mode]\" value=\"usecase\"  aria-describedby=\"tk-llm-help-mode-usecase\"><span>I have a use case<\/span><\/label><label><input type=\"radio\" name=\"llm[mode]\" value=\"compare\"  aria-describedby=\"tk-llm-help-mode-compare\"><span>Compare configurations<\/span><\/label><\/fieldset>\n <p class=\"tk-llm-mode-help\" id=\"tk-llm-help-mode-hardware\" data-modes=\"hardware\">Your hardware is known. The finder checks model configurations for this device.<\/p><p class=\"tk-llm-mode-help\" id=\"tk-llm-help-mode-model\" data-modes=\"model\">Your model is known. The finder checks which listed hardware can run it under these settings.<\/p><p class=\"tk-llm-mode-help\" id=\"tk-llm-help-mode-usecase\" data-modes=\"usecase\">Describe the task and requirements. The finder filters configurations using these inputs.<\/p><p class=\"tk-llm-mode-help\" id=\"tk-llm-help-mode-compare\" data-modes=\"compare\">Select two to four configurations for a side-by-side comparison.<\/p> <div class=\"tk-llm-compare\" id=\"tk-llm-compare\" data-modes=\"compare\"><h3>Select up to four configurations<\/h3><label class=\"tk-llm-compare-search\">Search configurations<input type=\"search\" name=\"llm[cq]\" value=\"\" maxlength=\"120\" autocomplete=\"off\" data-search-configurations aria-describedby=\"tk-llm-help-cq\"><\/label><small id=\"tk-llm-help-cq\">Search by model, quantisation or runtime. 1,092 configurations are available.<\/small><p class=\"tk-llm-selection-count\" role=\"status\" aria-live=\"polite\"><\/p><div class=\"tk-llm-checks tk-llm-configurations\" data-configurations><p class=\"tk-llm-compare-empty\">Nothing selected yet. Search above for a model to compare configurations.<\/p><\/div><\/div>\n <div class=\"tk-llm-grid tk-llm-primary\">\n <div data-modes=\"usecase\"><label for=\"tk-llm-field-task\">Task<select id=\"tk-llm-field-task\" name=\"llm[task]\" aria-describedby=\"tk-llm-help-task\"><option value=\"general\" selected>General<\/option><option value=\"rag\">RAG<\/option><option value=\"coding\">Coding<\/option><option value=\"agent\">Agents \/ tool use<\/option><option value=\"document\">Document analysis<\/option><\/select><\/label><small id=\"tk-llm-help-task\">The task sets the speed guide value for the shortlist. For agents, the finder also checks support for external tools. The task does not filter models by answer quality.<\/small><\/div>\n <div data-modes=\"hardware usecase compare\"><div class=\"tk-llm-catalog-field\"><label for=\"tk-llm-field-hardware\">Hardware<select id=\"tk-llm-field-hardware\" name=\"llm[hardware]\"><optgroup label=\"AMD\"><option value=\"amd-mi210\">AMD Instinct MI210<\/option><option value=\"amd-mi250x\">AMD Instinct MI250X module<\/option><option value=\"amd-mi300x\">AMD Instinct MI300X<\/option><option value=\"amd-mi325x\">AMD Instinct MI325X<\/option><option value=\"amd-mi350x\">AMD Instinct MI350X<\/option><option value=\"amd-r9700\">AMD Radeon AI PRO R9700<\/option><option value=\"amd-w7900\">AMD Radeon PRO W7900<\/option><option value=\"amd-rx7800xt\">AMD Radeon RX 7800 XT<\/option><option value=\"amd-rx7900xt\">AMD Radeon RX 7900 XT<\/option><option value=\"amd-rx7900xtx\">AMD Radeon RX 7900 XTX<\/option><option value=\"amd-rx9070\">AMD Radeon RX 9070<\/option><option value=\"amd-rx9070xt\">AMD Radeon RX 9070 XT<\/option><\/optgroup><optgroup label=\"Apple\"><option value=\"macbook-pro-14-m5-pro-64-gb\">MacBook Pro 14\u2033 \u00b7 M5 Pro \u00b7 64 GB<\/option><option value=\"macbook-pro-14-m5-32-gb\">MacBook Pro 14\u2033 \u00b7 M5 \u00b7 32 GB<\/option><option value=\"mac-mini-m1-16-gb\">Mac mini \u00b7 M1 \u00b7 16 GB<\/option><option value=\"mac-mini-m2-pro-32-gb\">Mac mini \u00b7 M2 Pro \u00b7 32 GB<\/option><option value=\"mac-mini-m2-24-gb\">Mac mini \u00b7 M2 \u00b7 24 GB<\/option><option value=\"mac-mini-m4-pro-48-gb\">Mac mini \u00b7 M4 Pro \u00b7 48 GB<\/option><option value=\"mac-mini-m4-pro-64-gb\">Mac mini \u00b7 M4 Pro \u00b7 64 GB<\/option><option value=\"mac-mini-m4-24-gb\">Mac mini \u00b7 M4 \u00b7 24 GB<\/option><option value=\"mac-pro-m2-ultra-192-gb\">Mac Pro \u00b7 M2 Ultra \u00b7 192 GB<\/option><option value=\"mac-studio-m1-max-64-gb\">Mac Studio \u00b7 M1 Max \u00b7 64 GB<\/option><option value=\"mac-studio-m1-ultra-128-gb\">Mac Studio \u00b7 M1 Ultra \u00b7 128 GB<\/option><option value=\"mac-studio-m2-max-96-gb\">Mac Studio \u00b7 M2 Max \u00b7 96 GB<\/option><option value=\"mac-studio-m3-ultra-256-gb\">Mac Studio \u00b7 M3 Ultra \u00b7 256 GB<\/option><option value=\"mac-studio-m4-max-128-gb\">Mac Studio \u00b7 M4 Max \u00b7 128 GB<\/option><option value=\"mac-studio-m5-max-128-gb\">Mac Studio \u00b7 M5 Max \u00b7 128 GB<\/option><option value=\"mac-studio-m5-ultra-512-gb\">Mac Studio \u00b7 M5 Ultra \u00b7 512 GB<\/option><\/optgroup><optgroup label=\"Intel\"><option value=\"intel-a750\">Intel Arc A750<\/option><option value=\"intel-a770-8\">Intel Arc A770 8GB<\/option><option value=\"intel-a770-16\">Intel Arc A770 16GB<\/option><option value=\"intel-b570\">Intel Arc B570<\/option><option value=\"intel-b580\">Intel Arc B580<\/option><option value=\"intel-pro-b50\">Intel Arc Pro B50<\/option><option value=\"intel-pro-b60\">Intel Arc Pro B60<\/option><\/optgroup><optgroup label=\"NVIDIA\"><option value=\"rtx-2060-6-gb\">GeForce RTX 2060 6 GB<\/option><option value=\"rtx-2060-12-gb\">GeForce RTX 2060 12 GB<\/option><option value=\"rtx-2060-super\">GeForce RTX 2060 SUPER<\/option><option value=\"rtx-2070-super\">GeForce RTX 2070 SUPER<\/option><option value=\"rtx-2080-ti\">GeForce RTX 2080 Ti<\/option><option value=\"rtx-3050-6-gb\">GeForce RTX 3050 6 GB<\/option><option value=\"rtx-3050-8-gb\">GeForce RTX 3050 8 GB<\/option><option value=\"rtx-3060-8-gb\">GeForce RTX 3060 8 GB<\/option><option value=\"rtx-3060-12-gb\">GeForce RTX 3060 12 GB<\/option><option value=\"rtx-3060-ti\">GeForce RTX 3060 Ti<\/option><option value=\"rtx-3070\">GeForce RTX 3070<\/option><option value=\"rtx-3070-ti\">GeForce RTX 3070 Ti<\/option><option value=\"rtx-3080-10-gb\">GeForce RTX 3080 10 GB<\/option><option value=\"rtx-3080-12-gb\">GeForce RTX 3080 12 GB<\/option><option value=\"rtx-3080-ti\">GeForce RTX 3080 Ti<\/option><option value=\"rtx-3090\">GeForce RTX 3090<\/option><option value=\"rtx-3090-ti\">GeForce RTX 3090 Ti<\/option><option value=\"rtx-4060\">GeForce RTX 4060<\/option><option value=\"rtx-4060-ti-8-gb\">GeForce RTX 4060 Ti 8 GB<\/option><option value=\"rtx-4060-ti-16-gb\">GeForce RTX 4060 Ti 16 GB<\/option><option value=\"rtx-4070\">GeForce RTX 4070<\/option><option value=\"rtx-4070-super\">GeForce RTX 4070 SUPER<\/option><option value=\"rtx-4070-ti\">GeForce RTX 4070 Ti<\/option><option value=\"rtx-4070-ti-super\">GeForce RTX 4070 Ti SUPER<\/option><option value=\"rtx-4080\">GeForce RTX 4080<\/option><option value=\"rtx-4080-super\">GeForce RTX 4080 SUPER<\/option><option value=\"rtx4090\" selected>GeForce RTX 4090<\/option><option value=\"rtx-5050\">GeForce RTX 5050<\/option><option value=\"rtx-5060\">GeForce RTX 5060<\/option><option value=\"rtx-5060-ti-8-gb\">GeForce RTX 5060 Ti 8 GB<\/option><option value=\"rtx-5060-ti-16-gb\">GeForce RTX 5060 Ti 16 GB<\/option><option value=\"rtx-5070\">GeForce RTX 5070<\/option><option value=\"rtx-5070-ti\">GeForce RTX 5070 Ti<\/option><option value=\"rtx-5080\">GeForce RTX 5080<\/option><option value=\"rtx-5090\">GeForce RTX 5090<\/option><option value=\"a100-40gb-pcie\">NVIDIA A100 40GB PCIe<\/option><option value=\"a100-80gb-pcie\">NVIDIA A100 80GB PCIe<\/option><option value=\"a100-80gb-sxm\">NVIDIA A100 80GB SXM<\/option><option value=\"b200-sxm\">NVIDIA B200 SXM<\/option><option value=\"b300-sxm\">NVIDIA B300 SXM<\/option><option value=\"h100-80gb-pcie\">NVIDIA H100 80GB PCIe<\/option><option value=\"h100-80gb-sxm\">NVIDIA H100 80GB SXM<\/option><option value=\"h200-nvl\">NVIDIA H200 NVL<\/option><option value=\"l40s\">NVIDIA L40S<\/option><option value=\"rtx-4500-ada-generation-desktop\">NVIDIA RTX 4500 Ada Generation Desktop<\/option><option value=\"rtx-5000-ada-generation-desktop\">NVIDIA RTX 5000 Ada Generation Desktop<\/option><option value=\"rtx-6000-ada-generation-desktop\">NVIDIA RTX 6000 Ada Generation Desktop<\/option><option value=\"rtx-a2000-desktop-6gb\">NVIDIA RTX A2000 Desktop 6GB<\/option><option value=\"rtx-a2000-desktop-12gb\">NVIDIA RTX A2000 Desktop 12GB<\/option><option value=\"rtx-a4000-desktop\">NVIDIA RTX A4000 Desktop<\/option><option value=\"rtx-a5000-desktop\">NVIDIA RTX A5000 Desktop<\/option><option value=\"rtx-a6000-desktop\">NVIDIA RTX A6000 Desktop<\/option><option value=\"rtx-pro-6000-blackwell-max-q-workstation-edition\">NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition<\/option><option value=\"rtx-pro-6000-blackwell-server-edition\">NVIDIA RTX PRO 6000 Blackwell Server Edition<\/option><option value=\"rtx-pro-6000-blackwell-workstation-edition\">NVIDIA RTX PRO 6000 Blackwell Workstation Edition<\/option><\/optgroup><optgroup label=\"Planning profiles\"><option value=\"cuda48\">Planning profile CUDA \u00b7 48 GiB<\/option><option value=\"cuda80\">Planning profile CUDA \u00b7 80 GiB<\/option><option value=\"unified64\">Planning profile Unified memory \u00b7 64 GiB<\/option><\/optgroup><\/select><\/label><\/div><\/div>\n <p class=\"tk-llm-small\" data-modes=\"usecase\"><a href=\"?llm%5Bmode%5D=usecase&amp;llm%5Btask%5D=general&amp;llm%5Blanguage%5D=any&amp;llm%5Bsort%5D=fit&amp;llm%5Brank%5D=v3&amp;llm%5Bparallel%5D=single&amp;llm%5Bkv_location%5D=gpu&amp;llm%5Bresult_filter%5D=all&amp;llm%5Bhardware%5D=cuda48&amp;llm%5Bmodel%5D=&amp;llm%5Bruntime%5D=&amp;llm%5Bquantization%5D=&amp;llm%5Bcontext%5D=8192&amp;llm%5Bconcurrency%5D=1&amp;llm%5Bgpus%5D=1&amp;llm%5Bprefill_batch%5D=512&amp;llm%5Bkv_bits%5D=16&amp;llm%5Breserve_percent%5D=10&amp;llm%5Bworkspace_low%5D=1&amp;llm%5Bworkspace_high%5D=3&amp;llm%5Bhost_ram%5D=64&amp;llm%5Bcustom_memory_gib%5D=0&amp;llm%5Bpage%5D=1&amp;llm%5Bper_page%5D=20&amp;llm%5Boffload%5D=&amp;llm%5Bcommercial%5D=&amp;llm%5Bindependent_language%5D=&amp;llm%5Bvision%5D=&amp;llm%5Btools%5D=&amp;llm%5Brun%5D=1\" data-llm-set='{\"hardware\":\"cuda48\",\"custom_memory_gib\":\"0\",\"page\":\"1\"}'>Hardware undecided: start with a 48 GiB planning profile<\/a><\/p>\n <div data-modes=\"model\"><div class=\"tk-llm-catalog-field\"><label for=\"tk-llm-field-model\">Model variant<select id=\"tk-llm-field-model\" name=\"llm[model]\" aria-describedby=\"tk-llm-help-model\"><option value=\"\" disabled selected>Choose a model<\/option><optgroup label=\"01-ai\"><option value=\"hf-37db082ea1c526d2a118\">Yi-Coder-1.5B<\/option><option value=\"hf-2e0e57705eb5bbea2eec\">Yi-6B-Chat<\/option><option value=\"hf-87c5a6189979b36345fb\">Yi-1.5-9B-Chat<\/option><\/optgroup><optgroup label=\"ai21labs\"><option value=\"hf-7f9ef06beae079f3f5b7\">AI21-Jamba2-3B<\/option><option value=\"hf-dd1acf08c386d10beda6\">AI21-Jamba-Reasoning-3B<\/option><option value=\"hf-05c8421eef06d1b5e546\">AI21-Jamba-Mini-1.7<\/option><option value=\"hf-748bc855846b2bf7b271\">AI21-Jamba2-Mini<\/option><option value=\"hf-5d6ec3b179dd4e45c6c9\">AI21-Jamba-Large-1.7<\/option><\/optgroup><optgroup label=\"allenai\"><option value=\"hf-88ef807f6bdbf125c921\">OLMo-1B-hf<\/option><option value=\"hf-6e28b8f7b2e3a9f1ad0f\">OLMo-2-0425-1B<\/option><option value=\"hf-3a3b1f46a985154c8d92\">OLMo-2-0425-1B-Instruct<\/option><option value=\"hf-5ebe300756bf12beb051\">OLMoE-1B-7B-0125-Instruct<\/option><option value=\"hf-95494f14f292b08a7173\">OLMoE-1B-7B-0924<\/option><option value=\"hf-c9470f603cdcb49cc227\">Olmo-3-7B-Instruct<\/option><option value=\"hf-c3361cf235565b390429\">Olmo-3-7B-Instruct-SFT<\/option><option value=\"hf-cfa132c83ed0d591f727\">Olmo-3-7B-Think<\/option><option value=\"hf-07d88b9c0c3df3669fda\">Olmo-3-1025-7B<\/option><option value=\"hf-cace834b195ae7732c91\">OLMo-2-1124-7B<\/option><option value=\"hf-ba49f1b850fbf7dc9027\">OLMo-2-1124-7B-Instruct<\/option><option value=\"hf-4f242b0fcb2f07c45a76\">Llama-3.1-Tulu-3-8B-SFT<\/option><option value=\"hf-4585478191fbb05233ca\">olmOCR-2-7B-1025<\/option><option value=\"hf-e1a4dd53ee86e7ff515e\">Olmo-3-32B-Think<\/option><option value=\"hf-aafb328ee9069fe09734\">OLMo-2-0325-32B-Instruct<\/option><\/optgroup><optgroup label=\"apple\"><option value=\"hf-69e731f9c96537ef7192\">OpenELM-1_1B-Instruct<\/option><option value=\"hf-107c84714fb665d098f5\">DiffuCoder-7B-cpGRPO<\/option><\/optgroup><optgroup label=\"arcee-ai\"><option value=\"hf-438f480f36605696b8df\">Arcee-VyLinh<\/option><option value=\"hf-eae69cecf79e00a27041\">AFM-4.5B<\/option><option value=\"hf-0217c92412f31433344f\">Trinity-Nano-Preview<\/option><option value=\"hf-405e645f41a92ea6eff9\">Llama-3.1-SuperNova-Lite<\/option><option value=\"hf-6afa1b95327b8ba1f0d4\">Virtuoso-Small-v2<\/option><option value=\"hf-ea704e4789b309f39ad2\">Virtuoso-Small<\/option><option value=\"hf-698b8a0a5e9e849f70f8\">Trinity-Mini<\/option><option value=\"hf-19284ba1dd82c15b7140\">Caller<\/option><option value=\"hf-5cff6fd400031a8a8d66\">Llama-3-SEC-Chat<\/option><option value=\"hf-7e809f2d64945cc46ac3\">Trinity-Large-Thinking<\/option><option value=\"hf-ff1d072829c69897f575\">Trinity-Large-TrueBase<\/option><\/optgroup><optgroup label=\"baidu\"><option value=\"hf-7054ac5fff4b82ef7ef3\">ERNIE-4.5-21B-A3B-Thinking<\/option><option value=\"hf-311dc891d92ee7b85e71\">ERNIE-4.5-VL-28B-A3B-Thinking<\/option><\/optgroup><optgroup label=\"ByteDance-Seed\"><option value=\"hf-d0c1b231a455063ea15b\">UI-TARS-2B-SFT<\/option><option value=\"hf-608c6655ef7855db6473\">Seed-Coder-8B-Instruct<\/option><option value=\"hf-48f99476230c532a1696\">UI-TARS-7B-DPO<\/option><option value=\"hf-51839fe4b87438fffb83\">UI-TARS-1.5-7B<\/option><option value=\"hf-1c63858b488cc335ee5e\">Seed-OSS-36B-Instruct<\/option><\/optgroup><optgroup label=\"CohereLabs\"><option value=\"hf-d0e50ed0044fa6e441f0\">North-Micro-Vision-Instruct<\/option><option value=\"hf-1c6197912678073d8182\">aya-expanse-8b<\/option><option value=\"hf-802673a10c89221effe7\">c4ai-command-r7b-12-2024<\/option><option value=\"hf-eb602270bd70597872bd\">aya-vision-8b<\/option><option value=\"hf-cfc0a4fde2f5ac235de2\">North-Mini-Code-1.0<\/option><option value=\"hf-f87f1180299af7d466b7\">aya-vision-32b<\/option><option value=\"hf-bba75bfcfbbe9cb552fe\">c4ai-command-r-v01<\/option><option value=\"hf-eceb232a9f38569dccd5\">c4ai-command-r-plus<\/option><option value=\"hf-2b7ab95855999b97c40d\">command-a-reasoning-08-2025<\/option><\/optgroup><optgroup label=\"DeepSeek\"><option value=\"deepseek-ai-deepseek-r1-distill-qwen-1.5b\">DeepSeek-R1-Distill-Qwen-1.5B<\/option><option value=\"hf-a5df55d251a644dc2c56\">deepseek-vl2-tiny<\/option><option value=\"hf-e56f1cd9ebc6e11b3d0d\">DeepSeek-Prover-V2-7B<\/option><option value=\"deepseek-ai-deepseek-r1-distill-qwen-7b\">DeepSeek-R1-Distill-Qwen-7B<\/option><option value=\"deepseek-ai-deepseek-r1-distill-llama-8b\">DeepSeek-R1-Distill-Llama-8B<\/option><option value=\"deepseek-ai-deepseek-r1-0528-qwen3-8b\">DeepSeek-R1-0528-Qwen3-8B<\/option><option value=\"deepseek-ai-deepseek-r1-distill-qwen-14b\">DeepSeek-R1-Distill-Qwen-14B<\/option><option value=\"hf-617407aadcc99469745d\">DeepSeek-Coder-V2-Lite-Instruct<\/option><option value=\"hf-d81be52bb954baf65b36\">DeepSeek-V2-Lite-Chat<\/option><option value=\"deepseek-ai-deepseek-r1-distill-qwen-32b\">DeepSeek-R1-Distill-Qwen-32B<\/option><option value=\"deepseek-ai-deepseek-r1-distill-llama-70b\">DeepSeek-R1-Distill-Llama-70B<\/option><option value=\"hf-77cd9239f83542017b45\">DeepSeek-Coder-V2-Instruct<\/option><option value=\"hf-1a0e66ba3eccfab66105\">DeepSeek-Coder-V2-Instruct-0724<\/option><option value=\"hf-efd5a5cfd797123c410b\">DeepSeek-V2-Chat-0628<\/option><option value=\"hf-366db638009a1d136aa6\">DeepSeek-V4-Flash<\/option><option value=\"hf-3c80a4a9b90355b60d5e\">DeepSeek-V4-Flash-0731<\/option><option value=\"hf-df3e05a17d77a6a8606d\">DeepSeek-R1<\/option><option value=\"deepseek-ai-deepseek-r1-0528\">DeepSeek-R1-0528<\/option><option value=\"hf-65845821798dff059aca\">DeepSeek-V3<\/option><option value=\"hf-d177e8bb6667959403d9\">DeepSeek-V3-0324<\/option><option value=\"deepseek-ai-deepseek-v3.2\">DeepSeek-V3.2<\/option><\/optgroup><optgroup label=\"Google\"><option value=\"hf-b2c212397053e81f1e7c\">gemma-3-270m<\/option><option value=\"hf-97fbad0733dd28ae94a4\">gemma-4-12B-it-assistant<\/option><option value=\"google-gemma-3-1b-it\">gemma-3-1b-it<\/option><option value=\"hf-33b217b2be53d9ca1491\">gemma-2-2b<\/option><option value=\"hf-87e059e18defb942b4aa\">gemma-2-2b-it<\/option><option value=\"google-gemma-3-4b-it\">gemma-3-4b-it<\/option><option value=\"hf-38a5b91cacac14c30067\">medgemma-1.5-4b-it<\/option><option value=\"hf-92b390914c2a5a955e00\">medgemma-4b-it<\/option><option value=\"google-gemma-4-e2b\">gemma-4-E2B<\/option><option value=\"hf-da07a1a9fdc37341a8df\">gemma-4-E2B-it<\/option><option value=\"google-gemma-4-e4b\">gemma-4-E4B<\/option><option value=\"hf-84b69a06fd1cc63ed207\">gemma-4-E4B-it<\/option><option value=\"hf-1f9bfb32cf6ad7589ecd\">gemma-2-9b-it<\/option><option value=\"hf-a6e0a6e0e1b8c0f85b2e\">gemma-4-12B<\/option><option value=\"hf-5a21e44b941135c37a5e\">gemma-4-12B-it<\/option><option value=\"google-gemma-3-12b-it\">gemma-3-12b-it<\/option><option value=\"hf-c6079e21213b744c341f\">gemma-4-26B-A4B-it<\/option><option value=\"hf-00de91117ce05b3beb35\">diffusiongemma-26B-A4B-it<\/option><option value=\"google-gemma-4-26b-a4b\">gemma-4-26B-A4B<\/option><option value=\"google-gemma-3-27b-it\">gemma-3-27b-it<\/option><option value=\"hf-410fa9a7b0347bd08808\">gemma-4-31B-it<\/option><option value=\"google-gemma-4-31b\">gemma-4-31B<\/option><\/optgroup><optgroup label=\"HuggingFaceTB\"><option value=\"hf-9eb9ff63418b4611143b\">SmolLM-135M<\/option><option value=\"hf-ffd1482026cf85dce293\">SmolLM-135M-Instruct<\/option><option value=\"hf-c2ff4a64443647050e33\">SmolLM2-135M-Instruct<\/option><option value=\"hf-9d56e5c6b27c2b8855a8\">SmolLM2-360M-Instruct<\/option><option value=\"hf-bc6d45ed056eb474edab\">SmolLM-1.7B<\/option><option value=\"hf-70bdbe36fddf1a4b3026\">SmolLM2-1.7B-Instruct<\/option><option value=\"hf-3468de1111a048730117\">SmolVLM-Instruct<\/option><option value=\"hf-3b44bd607b9faeb932e2\">SmolLM3-3B<\/option><\/optgroup><optgroup label=\"IBM Granite\"><option value=\"hf-6e65fe28ba043041905d\">granite-3.0-1b-a400m-instruct<\/option><option value=\"hf-29c86cb09861bf494322\">granite-3.1-1b-a400m-instruct<\/option><option value=\"hf-387387c151589697ddc2\">granite-3.1-2b-instruct<\/option><option value=\"ibm-granite-granite-3.3-2b-instruct\">granite-3.3-2b-instruct<\/option><option value=\"hf-83919a21841626f64fe2\">granite-vision-3.2-2b<\/option><option value=\"hf-c10e775ec982dcbdd254\">granite-4.0-micro<\/option><option value=\"hf-49a57769820009abff2d\">granite-4.1-3b<\/option><option value=\"hf-af0e45b984d673d185b2\">granite-4.2-3b<\/option><option value=\"hf-60f53a435a14ab590ceb\">granite-vision-4.1-4b<\/option><option value=\"hf-9498eaf25837b2a02b30\">granite-4.0-tiny-preview<\/option><option value=\"ibm-granite-granite-4.0-h-tiny\">granite-4.0-h-tiny<\/option><option value=\"hf-e0292d95cf7e429225bd\">granite-3.0-8b-instruct<\/option><option value=\"hf-25edf1e359965763b7bb\">granite-3.1-8b-instruct<\/option><option value=\"hf-79edb4b272cedac87b0d\">granite-3.2-8b-instruct<\/option><option value=\"ibm-granite-granite-3.3-8b-instruct\">granite-3.3-8b-instruct<\/option><option value=\"hf-28095b34d31fd98ee2a7\">granite-guardian-3.3-8b<\/option><option value=\"hf-c881546df5988ffd77ff\">granite-4.1-8b<\/option><option value=\"hf-d90c16215efd7e052fca\">granite-4.2-8b<\/option><option value=\"hf-49fa90df7a8e78612918\">granite-4.1-30b<\/option><option value=\"hf-f3730f12e4a86e5f0b39\">granite-4.2-30b<\/option><option value=\"ibm-granite-granite-4.0-h-small\">granite-4.0-h-small<\/option><\/optgroup><optgroup label=\"IFM\"><option value=\"hf-e39bfd4c77201628340f\">K2-Horizon-7B<\/option><option value=\"hf-0f613ec2a99e1b94d080\">K2-Horizon-32B<\/option><option value=\"hf-3605a0010f2d8ff7e0de\">K2-Horizon-MoVA-36B-A4B<\/option><option value=\"hf-1931052e2bd7b9b7dc16\">K2-Horizon-375B-A23B<\/option><option value=\"hf-b9556cf925d4d3c3badb\">K2-Chat<\/option><\/optgroup><optgroup label=\"inclusionAI\"><option value=\"hf-1a781ed8b63fc2b55b9d\">UI-Venus-2-9B<\/option><option value=\"hf-600a9cc57312b2c4117d\">DR-Venus-4B-SFT<\/option><option value=\"hf-791bae9c4b31435f9635\">Ling-3.0-tiny<\/option><option value=\"hf-9b0bd4149bbfb95cae92\">Ling-mini-2.0<\/option><option value=\"hf-ba1f4563d8b427ccf1a4\">LLaDA2.0-mini<\/option><option value=\"hf-6dc5af50c3955d0bbe8c\">LLaDA2.1-mini<\/option><option value=\"hf-724c2a784624afa22f73\">Ling-3.0-flash-VL<\/option><option value=\"hf-56d2fa9415bd05a4c52f\">Ling-3.0-flash<\/option><\/optgroup><optgroup label=\"internlm\"><option value=\"hf-6e6af22cb61f67ea5597\">Intern-S1-mini<\/option><option value=\"hf-996feeae6c9e47763667\">internlm3-8b-instruct<\/option><option value=\"hf-535b4389c5cbc61d6943\">Intern-S1<\/option><option value=\"hf-0e8ac03b63a25156289c\">Atria-Dawn-Preview<\/option><\/optgroup><optgroup label=\"JetBrains\"><option value=\"hf-b9282c99743c4636361a\">Mellum2-12B-A2.5B-Instruct<\/option><option value=\"hf-9c8220032f788d484c59\">Mellum2-12B-A2.5B-Thinking<\/option><\/optgroup><optgroup label=\"kakaocorp\"><option value=\"hf-0048fe63f4226c94ad5d\">kanana-2-3b-instruct<\/option><\/optgroup><optgroup label=\"LGAI-EXAONE\"><option value=\"hf-1abcc8273d8a69ecc9aa\">EXAONE-4.0-1.2B<\/option><option value=\"hf-ad6bef586969f54bd308\">EXAONE-3.5-2.4B-Instruct<\/option><option value=\"hf-7748da382a8126c659d6\">EXAONE-Deep-2.4B<\/option><option value=\"hf-fd5343254c29bf6000ff\">EXAONE-3.0-7.8B-Instruct<\/option><option value=\"hf-00e46d7a1ae474441790\">EXAONE-3.5-7.8B-Instruct<\/option><option value=\"hf-ad56d316a5c906b14c13\">EXAONE-Deep-7.8B<\/option><option value=\"hf-fa55aac6cc41aec7dedf\">EXAONE-3.5-32B-Instruct<\/option><option value=\"hf-a249fe04b62c403357e2\">EXAONE-4.0-32B<\/option><option value=\"hf-27d72721f494096860f5\">EXAONE-4.0.1-32B<\/option><option value=\"hf-b5ef5665f3daee9a987c\">EXAONE-4.5-33B<\/option><option value=\"hf-4281341036649e150470\">K-EXAONE-236B-A23B<\/option><option value=\"hf-7efc30874c88d5358526\">K-EXAONE-2.0-750B-A37B<\/option><\/optgroup><optgroup label=\"LiquidAI\"><option value=\"hf-36ee8f93037805eec3a8\">LFM2.5-230M<\/option><option value=\"hf-d4d1f15ac55a224b2c15\">LFM2.5-1.2B-Instruct-DSpark<\/option><option value=\"hf-9e67935db1fbc79ab351\">LFM2-350M-ENJP-MT<\/option><option value=\"hf-0b9844e5bfdffffaddd5\">LFM2.5-350M<\/option><option value=\"hf-1f750f907453d8806570\">LFM2.5-VL-450M<\/option><option value=\"hf-b31f7ae49db676278553\">LFM2-1.2B<\/option><option value=\"hf-8b614a1726aa3350cbb2\">LFM2.5-1.2B-Instruct<\/option><option value=\"hf-d07b2be385975a0e4b6a\">LFM2.5-1.2B-Thinking<\/option><option value=\"hf-68aa29618ae6b5810d85\">LFM2-VL-1.6B<\/option><option value=\"hf-dcbec49d68fbb4a8f0fd\">LFM2.5-VL-1.6B<\/option><option value=\"hf-01fab24b492e98aa1a54\">LFM2.5-2.6B<\/option><option value=\"hf-0029109f59f9df5e5a65\">LFM2.5-VL-3B<\/option><option value=\"hf-3e5f0b9c07a5ca455fb8\">LFM2-8B-A1B<\/option><option value=\"hf-0e4fd604629d7e82c222\">LFM2.5-8B-A1B<\/option><\/optgroup><optgroup label=\"llm-jp\"><option value=\"hf-bc41ad118882dc214604\">llm-jp-4-8b-thinking<\/option><option value=\"hf-f4a50b9746a12e442f55\">llm-jp-4-vl-9b<\/option><option value=\"hf-0f4076d0826a6afcaf10\">llm-jp-4-33b-thinking<\/option><\/optgroup><optgroup label=\"Meta\"><option value=\"meta-llama-llama-3.2-1b-instruct\">Llama-3.2-1B-Instruct<\/option><option value=\"hf-25ae3bbf3c04d2f7d88c\">Llama-3.2-3B<\/option><option value=\"meta-llama-llama-3.2-3b-instruct\">Llama-3.2-3B-Instruct<\/option><option value=\"hf-30842987926c0e3814b1\">Llama-3.1-8B<\/option><option value=\"meta-llama-llama-3.1-8b-instruct\">Llama-3.1-8B-Instruct<\/option><option value=\"hf-12a888a4670b0547c2b6\">Meta-Llama-3-8B<\/option><option value=\"meta-llama-meta-llama-3-8b-instruct\">Meta-Llama-3-8B-Instruct<\/option><option value=\"meta-llama-llama-3.2-11b-vision-instruct\">Llama-3.2-11B-Vision-Instruct<\/option><option value=\"meta-llama-llama-3.1-70b-instruct\">Llama-3.1-70B-Instruct<\/option><option value=\"llama3.3-70b\">Llama-3.3-70B-Instruct<\/option><option value=\"meta-llama-meta-llama-3-70b-instruct\">Meta-Llama-3-70B-Instruct<\/option><option value=\"meta-llama-llama-3.2-90b-vision-instruct\">Llama-3.2-90B-Vision-Instruct<\/option><option value=\"hf-e657cd04a0a86414a425\">Llama-4-Scout-17B-16E<\/option><option value=\"meta-llama-llama-4-scout-17b-16e-instruct\">Llama-4-Scout-17B-16E-Instruct<\/option><option value=\"meta-llama-llama-4-maverick-17b-128e-instruct\">Llama-4-Maverick-17B-128E-Instruct<\/option><option value=\"hf-e1c4bec96aa43db5179a\">Llama-3.1-405B<\/option><option value=\"meta-llama-llama-3.1-405b-instruct\">Llama-3.1-405B-Instruct<\/option><\/optgroup><optgroup label=\"Microsoft\"><option value=\"hf-ad016fd7a587644b8823\">Florence-2-large-ft<\/option><option value=\"hf-98cb65a37f11aecdc2b3\">Phi-3-mini-4k-instruct<\/option><option value=\"hf-584d61ba052d794a5103\">Phi-3-mini-128k-instruct<\/option><option value=\"hf-9023b5d1e3add666b615\">Phi-3.5-mini-instruct<\/option><option value=\"microsoft-phi-4-mini-instruct\">Phi-4-mini-instruct<\/option><option value=\"microsoft-phi-4-mini-reasoning\">Phi-4-mini-reasoning<\/option><option value=\"hf-75703919e4a2e3838223\">Phi-3-vision-128k-instruct<\/option><option value=\"hf-8ef193344ab7c69c6bf9\">Phi-3.5-vision-instruct<\/option><option value=\"hf-dd5a692eddc4519ee032\">Mage-VL<\/option><option value=\"hf-c9a82978e8ecbedd19b8\">Phi-4-multimodal-instruct<\/option><option value=\"hf-27038d02484458a1fee2\">Fara-7B<\/option><option value=\"microsoft-phi-4\">phi-4<\/option><option value=\"microsoft-phi-4-reasoning\">Phi-4-reasoning<\/option><option value=\"microsoft-phi-4-reasoning-plus\">Phi-4-reasoning-plus<\/option><option value=\"hf-5451770362650df286ee\">Phi-3.5-MoE-instruct<\/option><\/optgroup><optgroup label=\"MiniMax\"><option value=\"hf-28dcac50a4863fa25ba7\">MiniMax-M2<\/option><option value=\"hf-09094a77804a007f1fde\">MiniMax-M2.1<\/option><option value=\"hf-5062d41d746b534d3e50\">MiniMax-M2.7<\/option><option value=\"hf-7d0b67a51b9e6d617d11\">MiniMax-M2.5<\/option><option value=\"hf-33ba7fa10e263fc2af6a\">MiniMax-M3<\/option><\/optgroup><optgroup label=\"Mistral AI\"><option value=\"mistralai-ministral-3-3b-instruct-2512\">Ministral-3-3B-Instruct-2512<\/option><option value=\"mistralai-ministral-3-3b-reasoning-2512\">Ministral-3-3B-Reasoning-2512<\/option><option value=\"hf-9c744683d845426a9e66\">Voxtral-Mini-3B-2507<\/option><option value=\"hf-a639599a9e6a69f330dc\">Mistral-7B-Instruct-v0.2<\/option><option value=\"hf-42a64a672c6a4e0455e6\">Mistral-7B-Instruct-v0.3<\/option><option value=\"hf-b6206be5c26d77c6aeb4\">Mistral-7B-v0.3<\/option><option value=\"hf-16afab6e70d1b4b8d704\">Ministral-8B-Instruct-2410<\/option><option value=\"mistralai-ministral-3-8b-reasoning-2512\">Ministral-3-8B-Reasoning-2512<\/option><option value=\"mistralai-ministral-3-8b-instruct-2512\">Ministral-3-8B-Instruct-2512<\/option><option value=\"mistralai-mistral-nemo-instruct-2407\">Mistral-Nemo-Instruct-2407<\/option><option value=\"mistralai-ministral-3-14b-reasoning-2512\">Ministral-3-14B-Reasoning-2512<\/option><option value=\"mistralai-ministral-3-14b-instruct-2512\">Ministral-3-14B-Instruct-2512<\/option><option value=\"mistralai-codestral-22b-v0.1\">Codestral-22B-v0.1<\/option><option value=\"hf-0f4a84b4b394cb6d034a\">Devstral-Small-2507<\/option><option value=\"mistralai-magistral-small-2506\">Magistral-Small-2506<\/option><option value=\"mistralai-magistral-small-2509\">Magistral-Small-2509<\/option><option value=\"mistralai-mistral-small-3.1-24b-instruct-2503\">Mistral-Small-3.1-24B-Instruct-2503<\/option><option value=\"mistral-small-3.2\">Mistral-Small-3.2-24B-Instruct-2506<\/option><option value=\"mistralai-devstral-small-2-24b-instruct-2512\">Devstral-Small-2-24B-Instruct-2512<\/option><option value=\"hf-99012bba12773be387f1\">Voxtral-Small-24B-2507<\/option><option value=\"hf-a1138c0d5f61a9ba5361\">Mixtral-8x7B-Instruct-v0.1<\/option><option value=\"hf-b75173b780b3238c62ec\">Mixtral-8x7B-v0.1<\/option><option value=\"hf-c79387d109b51292e9be\">Mistral-Small-4-119B-2603<\/option><option value=\"hf-5c054278ba662af8a9be\">Mistral-Medium-3.5-128B<\/option><\/optgroup><optgroup label=\"Moonshot AI\"><option value=\"hf-53ad4b4966113729d3f8\">Moonlight-16B-A3B-Instruct<\/option><option value=\"hf-60ad859ce03327f2ac3d\">Kimi-Linear-48B-A3B-Instruct<\/option><option value=\"hf-c9f31b8dd14572f909fe\">Kimi-Dev-72B<\/option><option value=\"hf-7d737938e671c90fad24\">Kimi-K2-Instruct<\/option><option value=\"hf-ff7d6d6d44b092b6b0ed\">Kimi-K2-Thinking<\/option><option value=\"hf-41315cce6a165468fd04\">Kimi-K2-Instruct-0905<\/option><option value=\"hf-06434604eb1e2e26f931\">Kimi-K2.5<\/option><option value=\"hf-4a476e7ee2c72a9bb2a7\">Kimi-K2.6<\/option><option value=\"hf-5ea285ac7903cbc71f7e\">Kimi-K2.7-Code<\/option><\/optgroup><optgroup label=\"NousResearch\"><option value=\"hf-c0478dc17712311bd303\">Hermes-3-Llama-3.1-8B<\/option><option value=\"hf-d4db1c320e89f50852c2\">Meta-Llama-3-8B<\/option><option value=\"hf-66df882068c2c80b2087\">Meta-Llama-3-8B-Instruct<\/option><option value=\"hf-fdec12ff9ad25d049a1c\">Meta-Llama-3.1-8B<\/option><option value=\"hf-de6ecfffe8dbceb0faa5\">Meta-Llama-3.1-8B-Instruct<\/option><option value=\"hf-87aa3e56a754b9165826\">Meta-Llama-3-70B-Instruct<\/option><option value=\"hf-25ad12580193c1db6722\">Meta-Llama-3.1-70B-Instruct<\/option><\/optgroup><optgroup label=\"nvidia\"><option value=\"hf-8da8124444e98c42ed46\">Nemotron-H-8B-Reasoning-128K<\/option><option value=\"hf-de329d5d4bc9cf2f6b80\">Mistral-NeMo-Minitron-8B-Instruct<\/option><option value=\"hf-81c4eea80e54974655da\">NVIDIA-Nemotron-Nano-9B-v2<\/option><option value=\"hf-60883ee13cb7a3803862\">Nemotron-Cascade-14B-Thinking<\/option><option value=\"hf-dda6ef7654972439be0d\">Llama-3.1-Nemotron-70B-Instruct-HF<\/option><option value=\"hf-162aed1afbabbf2f900e\">Nemotron-Mini-4B-Instruct<\/option><\/optgroup><optgroup label=\"OpenAI\"><option value=\"openai-gpt-oss-20b\">gpt-oss-20b<\/option><option value=\"openai-gpt-oss-120b\">gpt-oss-120b<\/option><\/optgroup><optgroup label=\"openbmb\"><option value=\"hf-ab31824c5d59e70f9199\">MiniCPM5-1B<\/option><option value=\"hf-6fc7d720dc15cbfc5213\">MiniCPM-V-4.6<\/option><option value=\"hf-1d0213be4fb2cb82fbf5\">MiniCPM-V-4.6-Thinking<\/option><option value=\"hf-407d9206347b790dea39\">MiniCPM5-2B<\/option><option value=\"hf-743761581ad5d62a1a66\">MiniCPM-V-4<\/option><option value=\"hf-6eafd578009e59981901\">MiniCPM-V-2_6<\/option><option value=\"hf-f389ea593d832df53a69\">MiniCPM-Llama3-V-2_5<\/option><option value=\"hf-6bc0004acdc03e180527\">MiniCPM-o-2_6<\/option><option value=\"hf-a4f8025fb25c9d479008\">MiniCPM-V-4_5<\/option><option value=\"hf-b0480082235ce9823738\">MiniCPM-o-4_5<\/option><option value=\"hf-04cba501fb1f65128653\">MiniCPM-2B-dpo-bf16<\/option><option value=\"hf-599a7973630816103ff8\">MiniCPM3-4B<\/option><\/optgroup><optgroup label=\"openGPT-X\"><option value=\"hf-c096acedec576be9b274\">Teuken-7B-instruct-commercial-v0.4<\/option><option value=\"hf-b26c26db043a08bc71d7\">Teuken-7B-instruct-research-v0.4<\/option><\/optgroup><optgroup label=\"PleIAs\"><option value=\"hf-0738a8f59400afd5fc9d\">Baguettotron<\/option><option value=\"hf-10398e713493cca9d59b\">Pleias-RAG-350M<\/option><option value=\"hf-e31a208077fdf819f979\">Pleias-RAG-1B<\/option><\/optgroup><optgroup label=\"Qwen\"><option value=\"hf-d67ad4b44fe59dbd709c\">Qwen2.5-0.5B-Instruct<\/option><option value=\"qwen-qwen3-0.6b\">Qwen3-0.6B<\/option><option value=\"qwen-qwen3.5-0.8b\">Qwen3.5-0.8B<\/option><option value=\"hf-d91013282799672f5efe\">Qwen2.5-1.5B-Instruct<\/option><option value=\"qwen-qwen3-1.7b\">Qwen3-1.7B<\/option><option value=\"hf-5d05871d73da8d0b22f1\">Qwen3-VL-2B-Instruct<\/option><option value=\"hf-f10c85ac58beebd8dd45\">Qwen2-VL-2B-Instruct<\/option><option value=\"qwen-qwen3.5-2b\">Qwen3.5-2B<\/option><option value=\"hf-8c686effc3f68aad9b66\">Qwen2.5-3B-Instruct<\/option><option value=\"hf-92209a303d9d39a81a40\">Qwen2.5-VL-3B-Instruct<\/option><option value=\"qwen-qwen3-4b\">Qwen3-4B<\/option><option value=\"hf-05c0d3c7b35d797beea2\">Qwen3-4B-Instruct-2507<\/option><option value=\"hf-89942e47558e3a1f7ee9\">Qwen3-VL-4B-Instruct<\/option><option value=\"qwen-qwen3.5-4b\">Qwen3.5-4B<\/option><option value=\"hf-6a2bc752073cb97b5afb\">Qwen2.5-7B-Instruct<\/option><option value=\"hf-8deac41fee54fdddd068\">Qwen2.5-Coder-7B-Instruct<\/option><option value=\"qwen3-8b\">Qwen3-8B<\/option><option value=\"hf-cf559c8e863a21bf655f\">Qwen2.5-VL-7B-Instruct<\/option><option value=\"hf-749554b47437049d2ce4\">Qwen3-VL-8B-Instruct<\/option><option value=\"hf-12da7972176aebe73747\">Qwen3-VL-8B-Thinking<\/option><option value=\"qwen-qwen3.5-9b\">Qwen3.5-9B<\/option><option value=\"hf-5a23080f473c0891aaa8\">Qwen2.5-Omni-7B<\/option><option value=\"qwen3-14b\">Qwen3-14B<\/option><option value=\"hf-ff14dc6607c37448fad2\">Qwen2.5-14B-Instruct<\/option><option value=\"qwen-qwen3.5-27b\">Qwen3.5-27B<\/option><option value=\"qwen-qwen3.6-27b\">Qwen3.6-27B<\/option><option value=\"hf-401ffb2c754c3187f251\">Qwen3.8-27B<\/option><option value=\"qwen3-30b-a3b\">Qwen3-30B-A3B<\/option><option value=\"qwen-qwen3-30b-a3b-instruct-2507\">Qwen3-30B-A3B-Instruct-2507<\/option><option value=\"qwen-qwen3-coder-30b-a3b-instruct\">Qwen3-Coder-30B-A3B-Instruct<\/option><option value=\"hf-d23f99f8ee65669a557f\">Qwen3-VL-30B-A3B-Instruct<\/option><option value=\"hf-46011e1891325906a736\">Qwen3-VL-30B-A3B-Thinking<\/option><option value=\"qwen3-32b\">Qwen3-32B<\/option><option value=\"hf-42a684bad2fa4666a476\">Qwen3-Omni-30B-A3B-Instruct<\/option><option value=\"qwen-qwen3.5-35b-a3b\">Qwen3.5-35B-A3B<\/option><option value=\"qwen-qwen3.6-35b-a3b\">Qwen3.6-35B-A3B<\/option><option value=\"qwen-qwen3-coder-next\">Qwen3-Coder-Next<\/option><option value=\"qwen-qwen3.5-122b-a10b\">Qwen3.5-122B-A10B<\/option><option value=\"hf-c7963fe4056b14698818\">Qwen3.8-Flash-Next<\/option><option value=\"qwen-qwen3-235b-a22b-instruct-2507\">Qwen3-235B-A22B-Instruct-2507<\/option><option value=\"hf-2ef24b9cdfd62ca8db5b\">Qwen3-VL-235B-A22B-Instruct<\/option><option value=\"hf-dfa730408a21751b4ac6\">Qwen3-VL-235B-A22B-Thinking<\/option><option value=\"qwen-qwen3.5-397b-a17b\">Qwen3.5-397B-A17B<\/option><\/optgroup><optgroup label=\"sarvamai\"><option value=\"hf-da139269abf91014dc2f\">sarvam-1<\/option><option value=\"hf-273034872c256aad048e\">sarvam-translate<\/option><option value=\"hf-98b1495a24e61f094ce6\">sarvam-30b<\/option><option value=\"hf-481fb0bf91270b1974ba\">sarvam-105b<\/option><\/optgroup><optgroup label=\"ServiceNow-AI\"><option value=\"hf-4511f4752dc1fb105014\">Apriel-1.5-15b-Thinker<\/option><option value=\"hf-b6690502c771e41d9bb7\">Apriel-1.6-15b-Thinker<\/option><option value=\"hf-dc691ee3c99c13f1992d\">Apriel-Nemotron-15b-Thinker<\/option><\/optgroup><optgroup label=\"Skywork\"><option value=\"hf-40a8f9deea40652dbc78\">Skywork-OR1-7B<\/option><option value=\"hf-6c393108b88242c81656\">Skywork-OR1-Math-7B<\/option><option value=\"hf-bcbca74b69d6af371aaf\">Skywork-o1-Open-Llama-3.1-8B<\/option><option value=\"hf-93971dc75498f649d575\">Skywork-OR1-32B<\/option><option value=\"hf-e6c1bf591795a125c23d\">Skywork-R1V3-38B<\/option><\/optgroup><optgroup label=\"stabilityai\"><option value=\"hf-42e6399a2e6756ac4254\">stablelm-2-1_6b<\/option><option value=\"hf-18f2927bfd8d45d7cb96\">stablelm-2-1_6b-chat<\/option><option value=\"hf-36f7aecac56d341df362\">stablelm-2-zephyr-1_6b<\/option><option value=\"hf-a79ca46ae050d4752b9a\">stable-code-3b<\/option><option value=\"hf-95ab572eef94a4e3d188\">stable-code-instruct-3b<\/option><option value=\"hf-ccd2b7ec961e0940859d\">stablelm-2-12b-chat<\/option><\/optgroup><optgroup label=\"stepfun-ai\"><option value=\"hf-ff32f912cedc3812a3fa\">GOT-OCR2_0<\/option><option value=\"hf-980a2bed826bc87ef1e3\">Step-3.5-Flash<\/option><\/optgroup><optgroup label=\"swiss-ai\"><option value=\"hf-dc535ea2523843ac3d42\">Apertus-8B-2509<\/option><option value=\"hf-7e121d1a89527d463f0e\">Apertus-8B-Instruct-2509<\/option><option value=\"hf-da4d42ce792c483b3459\">Apertus-70B-Instruct-2509<\/option><\/optgroup><optgroup label=\"tencent\"><option value=\"hf-4d85f80b37cef68c74be\">Hunyuan-0.5B-Instruct<\/option><option value=\"hf-4b15575beeaeeb6fbaaf\">Hunyuan-1.8B-Instruct<\/option><option value=\"hf-93248927c814960a7a98\">Youtu-LLM-2B<\/option><option value=\"hf-9b0b258f6d3de9d6a03f\">HY-MT1.5-1.8B<\/option><option value=\"hf-a183600b18cb048e571f\">Hy-MT2-1.8B<\/option><option value=\"hf-5a21276bb142428e09cf\">Hunyuan-4B-Instruct<\/option><option value=\"hf-a64b33d58ee3526b4530\">Youtu-VL-4B-Instruct<\/option><option value=\"hf-9b0bb5ef90fc0b35cf21\">Hunyuan-7B-Instruct<\/option><option value=\"hf-2f352df9e6e46ae8167e\">Hy-MT2-7B<\/option><option value=\"hf-806553a2e30f8f369e7d\">Hunyuan-MT-7B<\/option><option value=\"hf-8e8cb3ea8406d1a2a9f8\">UI-Mate-9B<\/option><option value=\"hf-622502aed30ce0f06dbd\">UI-Mate-27B<\/option><option value=\"hf-144fb70b6e907066a50b\">Hy-MT2-30B-A3B<\/option><option value=\"hf-0e1dd02588b337321300\">Hunyuan-A13B-Instruct<\/option><option value=\"hf-d23e83fbb1d3aabc9fb3\">Hy3<\/option><option value=\"hf-10d7a91a76d7f8630949\">Hy3-preview<\/option><option value=\"hf-383f40728e3e52598b8a\">Hy4-preview<\/option><\/optgroup><optgroup label=\"tiiuae\"><option value=\"hf-6c47c60a3c0c3d5e6bed\">Falcon-H1-Tiny-90M-Instruct<\/option><option value=\"hf-c7bc5b4247085557331b\">Falcon-H1-Tiny-90M-Instruct-Curriculum<\/option><option value=\"hf-0a98b4cdf33ce105fac6\">Falcon-H1-Tiny-90M-Instruct-Curriculum-pre-DPO<\/option><option value=\"hf-f7e6e42cafee415e0349\">Falcon-H1-Tiny-90M-Instruct-pre-DPO<\/option><option value=\"hf-0f47cdab70ff5e25adb1\">Falcon-H1-Tiny-Multilingual-100M-Instruct<\/option><option value=\"hf-65a37087662222db7dbf\">Falcon-H1-0.5B-Instruct<\/option><option value=\"hf-6ca266629ba6b4ba7d32\">Falcon-E-1B-Instruct<\/option><option value=\"hf-442b00bf76ad4b542882\">Falcon-H1-Tiny-R-0.6B<\/option><option value=\"hf-f249fa5b234ebe20a436\">Falcon-H1-1.5B-Deep-Instruct<\/option><option value=\"hf-6981d1bc8b26fc55360a\">Falcon3-1B-Instruct<\/option><option value=\"hf-c71a817db889d54a53d4\">Falcon3-3B-Instruct<\/option><option value=\"hf-cdde5e75a4d4e1f04fe2\">falcon-mamba-7b<\/option><option value=\"hf-3f7bb78ae5d8af224d59\">falcon-mamba-7b-instruct<\/option><option value=\"hf-0c9cbcea0b9457295657\">Falcon3-7B-Instruct<\/option><option value=\"hf-354aa8441c2022bc4cab\">Falcon-H1-7B-Instruct<\/option><option value=\"hf-beb1cc4528f08c31ddf6\">Falcon3-10B-Instruct<\/option><\/optgroup><optgroup label=\"trillionlabs\"><option value=\"hf-62d148e96a2b74bec2bd\">Trillion-7B-preview<\/option><\/optgroup><optgroup label=\"upstage\"><option value=\"hf-941c91b5ec67b2b0a3cc\">Solar-Open-100B<\/option><\/optgroup><optgroup label=\"utter-project\"><option value=\"hf-c470575b792258ac2fc6\">EuroLLM-9B-Instruct<\/option><option value=\"hf-afd054c9b32e25f7a8ea\">EuroLLM-22B-Instruct-2512<\/option><\/optgroup><optgroup label=\"xai-org\"><option value=\"hf-0ed407b8c58398e6058c\">grok-2<\/option><\/optgroup><optgroup label=\"XiaomiMiMo\"><option value=\"hf-89ad3c72262d7d687df9\">MiMo-7B-RL<\/option><option value=\"hf-0fd71097dac2bd1ce9fa\">MiMo-VL-7B-RL<\/option><option value=\"hf-d7b8d34acc20505ad4c7\">MiMo-VL-7B-RL-2508<\/option><option value=\"hf-fb78bf48aac060693b06\">MiMo-VL-7B-SFT-2508<\/option><option value=\"hf-52763777e90f394b86b0\">MiMo-V2-Flash<\/option><option value=\"hf-00d1c6f432abd59f2a65\">MiMo-V2.5<\/option><\/optgroup><optgroup label=\"Z.ai\"><option value=\"hf-7e5dab001109ac169a33\">glm-edge-1.5b-chat<\/option><option value=\"hf-f2f6479a35590491e539\">glm-edge-4b-chat<\/option><option value=\"hf-f47434d72641f3483beb\">glm-4-9b-chat<\/option><option value=\"hf-a5c1565204aa1c2c5479\">glm-4-9b-chat-1m<\/option><option value=\"hf-fe75638340b39911fd9d\">GLM-4.1V-9B-Thinking<\/option><option value=\"hf-fcc3a1d31217f6f25c1f\">GLM-4.6V-Flash<\/option><option value=\"hf-557ddca7dae1d8c3e774\">GLM-4.7-Flash<\/option><option value=\"hf-950a323d193b2e37a3df\">GLM-4.5V<\/option><option value=\"hf-c1b61e1927c4ffb8abfb\">GLM-4.6V<\/option><option value=\"hf-0667055611ec08faaef7\">GLM-4.5-Air<\/option><option value=\"hf-25c23c4660a842609d04\">GLM-5.3-Flash<\/option><option value=\"hf-a6956a8ab574cb413e31\">GLM-4.7<\/option><option value=\"hf-15e1e78a69e85ede8125\">GLM-5.2<\/option><option value=\"hf-c49a2cc402b913c10b6a\">GLM-5.3<\/option><option value=\"hf-8920bd9ab39327e43076\">GLM-5.1<\/option><\/optgroup><\/select><\/label><small id=\"tk-llm-help-model\">Another 681 models are in the catalogue and are not calculated: 402 without a quantized edition at an allow-listed account, 254 no reason recorded, 20 not a language model choice, 5 format edition of another model. <a href=\"#tk-llm-catalog\">Open the model overview<\/a><\/small><\/div><\/div>\n <div class=\"tk-llm-preset\"><label class=\"tk-llm-preset-select\" hidden>Context per request (tokens)<select data-preset=\"context\" aria-describedby=\"tk-llm-help-context\"><option value=\"2048\">2,048 tokens<\/option><option value=\"4096\">4,096 tokens<\/option><option value=\"8192\" selected>8,192 tokens<\/option><option value=\"16384\">16,384 tokens<\/option><option value=\"32768\">32,768 tokens<\/option><option value=\"65536\">65,536 tokens<\/option><option value=\"131072\">131,072 tokens<\/option><option value=\"262144\">262,144 tokens<\/option><option value=\"custom\">Custom value<\/option><\/select><\/label><label class=\"tk-llm-custom\">Context per request (tokens)<input type=\"number\" required name=\"llm[context]\" min=\"256\" max=\"1048576\" value=\"8192\" aria-describedby=\"tk-llm-help-context\"><\/label><small id=\"tk-llm-help-context\">Tokens per request, including input, history and planned output.<\/small><\/div> <div class=\"tk-llm-preset\"><label class=\"tk-llm-preset-select\" hidden>Concurrent requests<select data-preset=\"concurrency\" aria-describedby=\"tk-llm-help-concurrency\"><option value=\"1\" selected>1<\/option><option value=\"2\">2<\/option><option value=\"4\">4<\/option><option value=\"8\">8<\/option><option value=\"16\">16<\/option><option value=\"32\">32<\/option><option value=\"64\">64<\/option><option value=\"custom\">Custom value<\/option><\/select><\/label><label class=\"tk-llm-custom\">Concurrent requests<input type=\"number\" required name=\"llm[concurrency]\" min=\"1\" max=\"128\" value=\"1\" aria-describedby=\"tk-llm-help-concurrency\"><\/label><small id=\"tk-llm-help-concurrency\">Number of requests processed at the same time.<\/small><\/div> <div data-modes=\"usecase\"><label for=\"tk-llm-field-language\">Application language<select id=\"tk-llm-field-language\" name=\"llm[language]\"><option value=\"any\" selected>Any<\/option><option value=\"de\">German<\/option><option value=\"en\">English<\/option><option value=\"fr\">French<\/option><\/select><\/label><\/div>\n <\/div><div class=\"tk-llm-checks\" data-modes=\"usecase\"><label><input type=\"checkbox\" name=\"llm[vision]\" value=\"1\" >Process images \/ scanned documents<\/label><label><input type=\"checkbox\" name=\"llm[tools]\" value=\"1\" >Require tool use<\/label><\/div>\n <details class=\"tk-llm-advanced\"><summary>Memory and serving assumptions<\/summary><div class=\"tk-llm-checks\"><div class=\"tk-llm-check\" ><label><input type=\"checkbox\" name=\"llm[commercial]\" value=\"1\"  aria-describedby=\"tk-llm-help-commercial\"> <span>For business use<\/span><\/label><small class=\"tk-llm-field-help\" id=\"tk-llm-help-commercial\">Missing license information is marked as unresolved.<\/small><\/div><div class=\"tk-llm-check\" data-modes=\"usecase\"><label><input type=\"checkbox\" name=\"llm[independent_language]\" value=\"1\"  aria-describedby=\"tk-llm-help-independent_language\"> <span>Independent tests for the selected language<\/span><\/label><small class=\"tk-llm-field-help\" id=\"tk-llm-help-independent_language\">Select a language. Without an independent test, suitability remains unverified.<\/small><\/div><div class=\"tk-llm-check\" ><label><input type=\"checkbox\" name=\"llm[offload]\" value=\"1\" > <span>Consider CPU offload<\/span><\/label><\/div><\/div>\n <div class=\"tk-llm-grid\"><div data-modes=\"hardware model usecase\"><label for=\"tk-llm-field-runtime\">Runtime<select id=\"tk-llm-field-runtime\" name=\"llm[runtime]\" aria-describedby=\"tk-llm-help-runtime\"><option value=\"\" selected>All<\/option><option value=\"llamacpp\">llama.cpp<\/option><option value=\"mlx\">MLX<\/option><option value=\"vllm\">vLLM<\/option><\/select><\/label><small id=\"tk-llm-help-runtime\">The runtime executes the model. Its version, format and hardware support must align.<\/small><\/div><div data-modes=\"hardware model usecase\"><label for=\"tk-llm-field-quantization\">Quantization<select id=\"tk-llm-field-quantization\" name=\"llm[quantization]\" aria-describedby=\"tk-llm-help-quantization\"><option value=\"\" selected>All<\/option><option value=\"awq_4bit\">AWQ 4bit<\/option><option value=\"bf16\">BF16<\/option><option value=\"fp8\">FP8<\/option><option value=\"gptq_4bit\">GPTQ 4bit<\/option><option value=\"mlx_4bit\">MLX 4bit<\/option><option value=\"mlx_8bit\">MLX 8bit<\/option><option value=\"nvfp4\">NVFP4<\/option><option value=\"q4_k_m\">Q4_K_M<\/option><option value=\"q8_0\">Q8_0<\/option><\/select><\/label><small id=\"tk-llm-help-quantization\">Quantization sets how many bits store each model value. It affects memory use and may affect quality.<\/small><\/div>\n <label for=\"tk-llm-field-parallel\">Distribution<select id=\"tk-llm-field-parallel\" name=\"llm[parallel]\"><option value=\"single\" selected>Single device<\/option><option value=\"layer\">Layer split (estimated)<\/option><option value=\"tensor\">Tensor split (estimated)<\/option><\/select><\/label><label for=\"tk-llm-field-kv_bits\">KV cache: bits per value<select id=\"tk-llm-field-kv_bits\" name=\"llm[kv_bits]\"><option value=\"16\" selected>16<\/option><option value=\"8\">8<\/option><\/select><\/label><label for=\"tk-llm-field-kv_location\">KV cache: location<select id=\"tk-llm-field-kv_location\" name=\"llm[kv_location]\"><option value=\"gpu\" selected>GPU \/ unified memory<\/option><option value=\"cpu\">CPU RAM<\/option><\/select><\/label> <label >Number of identical devices<input type=\"number\" required name=\"llm[gpus]\" min=\"1\" max=\"8\" step=\"1\" value=\"1\"><\/label><label data-modes=\"hardware usecase compare\">Custom memory planning per device (GiB, 0 = hardware value)<input type=\"number\" required name=\"llm[custom_memory_gib]\" min=\"0\" max=\"8192\" step=\"0.1\" value=\"0\"><\/label><label >Reserve per device (%)<input type=\"number\" required name=\"llm[reserve_percent]\" min=\"0\" max=\"50\" step=\"1\" value=\"10\"><\/label><label >Additional memory (workspace), minimum per device in GiB<input type=\"number\" required name=\"llm[workspace_low]\" min=\"0\" max=\"64\" step=\"0.1\" value=\"1\" aria-describedby=\"tk-llm-help-assumptions\"><\/label><label >Additional memory (workspace), maximum per device in GiB<input type=\"number\" required name=\"llm[workspace_high]\" min=\"0\" max=\"128\" step=\"0.1\" value=\"3\" aria-describedby=\"tk-llm-help-assumptions\"><\/label><label >Host RAM (GiB)<input type=\"number\" required name=\"llm[host_ram]\" min=\"1\" max=\"4096\" step=\"1\" value=\"64\"><\/label><\/div><p class=\"tk-llm-small\" id=\"tk-llm-help-assumptions\">Workspace is additional memory for intermediate results and the runtime. The range is an estimate. It depends on how many tokens are processed together and which runtime you use. Image processing may require additional memory.<\/p><p class=\"tk-llm-small\">The KV cache stores intermediate request states. Context length and concurrent requests increase its memory use.<\/p><p class=\"tk-llm-small\">GPU memory and CPU RAM may be used together, depending on the configuration. Actual allocation may differ.<\/p><\/details>\n <div class=\"tk-llm-actions\"><label for=\"tk-llm-field-result_filter\">Show results<select id=\"tk-llm-field-result_filter\" name=\"llm[result_filter]\"><option value=\"all\" selected>All<\/option><option value=\"memory_fit\">Enough memory<\/option><option value=\"viable\">All recorded requirements met<\/option><option value=\"unknown\">Review needed<\/option><\/select><\/label><label for=\"tk-llm-field-sort\">Sort by<select id=\"tk-llm-field-sort\" name=\"llm[sort]\"><option value=\"fit\" selected>First test candidate first<\/option><option value=\"memory\">Lowest memory requirement first<\/option><option value=\"speed\">Highest estimated speed first<\/option><option value=\"name\">Name<\/option><\/select><\/label><button type=\"submit\">Check configurations<\/button><\/div>\n <input type=\"hidden\" name=\"llm[page]\" value=\"1\"><input type=\"hidden\" name=\"llm[per_page]\" value=\"20\"><input type=\"hidden\" name=\"llm[run]\" value=\"1\">\n <\/form><p class=\"tk-llm-message\" role=\"status\" aria-live=\"polite\" aria-atomic=\"true\"><\/p>\n <div class=\"tk-llm-output\"><h3 id=\"tk-llm-results-title\" tabindex=\"-1\" aria-describedby=\"tk-llm-results-status\">Results <span class=\"tk-llm-count\">1092<\/span><\/h3>\n<p id=\"tk-llm-results-status\" class=\"tk-llm-small\">20 of 1,092 results, page 1 of 55.<\/p>\n<p class=\"tk-llm-small tk-llm-result-summary\">Calculated memory requirement: Fits the memory estimate: 83 \u00b7 Fits with little headroom: 18 \u00b7 At the memory limit: 5 \u00b7 Memory requirement incomplete: 686 \u00b7 Needs more memory: 300 \u00b7 Data snapshot <time datetime=\"2026-09-19\">Sep 19, 2026<\/time><\/p>\n<div class=\"tk-llm-legend\"><p class=\"tk-llm-why-title\">What the symbols mean<\/p><ul class=\"tk-llm-legend-list\" role=\"list\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span><span><b>clear<\/b>: Memory is sufficient by calculation, or the figure beside it is measured.<\/span><\/li><li class=\"is-warn\"><span aria-hidden=\"true\">!<\/span><span><b>with a caveat<\/b>: Tight, at the memory limit, or only with CPU offload. The sentence beside it names the case.<\/span><\/li><li class=\"is-fail\"><span aria-hidden=\"true\">\u2715<\/span><span><b>not possible<\/b>: Memory is missing, or runtime and hardware do not fit together.<\/span><\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span><span><b>open<\/b>: A value is missing from the data snapshot, so it cannot be calculated.<\/span><\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span><span><b>estimated<\/b>: Calculated, not measured.<\/span><\/li><li class=\"is-info\"><span aria-hidden=\"true\">i<\/span><span><b>measured<\/b>: Measured, source and origin are stated in the details.<\/span><\/li><\/ul><\/div><section class=\"tk-llm-recommendations\" aria-labelledby=\"tk-llm-shortlist-title\"><h4 id=\"tk-llm-shortlist-title\">Models for an initial test<\/h4><p class=\"tk-llm-small\">Guide value for this task: 10 tokens per second. That is an assumption by torck, no sourced threshold exists for it.<\/p><ul class=\"tk-llm-recommendation-grid\" role=\"list\">\n<li><article class=\"tk-llm-card tk-llm-recommendation is-first\" aria-labelledby=\"tk-llm-row-f29fc6161377-s-title\"><p class=\"tk-llm-recommendation-label\">First test candidate<\/p><h5 id=\"tk-llm-row-f29fc6161377-s-title\">Devstral-Small-2-24B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Nov 28, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Devstral-Small-2-24B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Devstral-Small-2-24B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><p class=\"tk-llm-selection-reason\">Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.<\/p><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Devstral-Small-2-24B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><p class=\"tk-llm-row-link\"><a href=\"#tk-llm-row-f29fc6161377\">Why this assessment: calculation and sources<\/a><\/p><\/article><\/li><li><article class=\"tk-llm-card tk-llm-recommendation\" aria-labelledby=\"tk-llm-row-f1609d8fb204-s-title\"><p class=\"tk-llm-recommendation-label\">Next place in the ranking<\/p><h5 id=\"tk-llm-row-f1609d8fb204-s-title\">Magistral-Small-2509<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Sep 12, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2509\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Magistral-Small-2509<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><p class=\"tk-llm-selection-reason\">Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.<\/p><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2509\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Magistral-Small-2509 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><p class=\"tk-llm-row-link\"><a href=\"#tk-llm-row-f1609d8fb204\">Why this assessment: calculation and sources<\/a><\/p><\/article><\/li><li><article class=\"tk-llm-card tk-llm-recommendation\" aria-labelledby=\"tk-llm-row-50468b18e01e-s-title\"><p class=\"tk-llm-recommendation-label\">Next place in the ranking<\/p><h5 id=\"tk-llm-row-50468b18e01e-s-title\">Mistral-Small-3.2-24B-Instruct-2506<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Jun 19, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.2-24B-Instruct-2506\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Mistral-Small-3.2-24B-Instruct-2506<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-info\"><span aria-hidden=\"true\">i<\/span> Suitability: vendor data only<\/li><\/ul><p class=\"tk-llm-selection-reason\">Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.<\/p><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><div class=\"tk-llm-evaluation\"><span class=\"tk-llm-quality-state\">Vendor claim<\/span><p>MMLU Pro: 69.1%<\/p><small>Figure for the model variant. No measurement is available for this quantization.<\/small><small>Vendor-reported with five examples and chain of thought. Not independently reproduced.<\/small><small><a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.2-24B-Instruct-2506\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face Hub \u00b7 Jun 20, 2025<\/a><\/small><\/div><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><p class=\"tk-llm-row-link\"><a href=\"#tk-llm-row-50468b18e01e\">Why this assessment: calculation and sources<\/a><\/p><\/article><\/li><\/ul>\n<p class=\"tk-llm-shortlist-rank\">The whole list follows one ranking. Configurations with verified data come first. Then it counts whether the estimated minimum speed at the selected context reaches the guide value for your task. After that the larger size class wins, within one size class the model whose repository was created later, then the higher parameter count, finally memory headroom and verified runtime data. The order says nothing about answer quality.<\/p>\n<\/section><details class=\"tk-llm-decision\"><summary>Decision brief and test plan<\/summary><div class=\"tk-llm-decision-report\" lang=\"en\" data-snapshot=\"torck-sync-a6d645b6456edf69b9b75e2a\"><h4>Decision brief and test plan<\/h4><p>Data snapshot: Sep 19, 2026<\/p><p>Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.<\/p><p>This selection provides candidates for a practical test. Memory and speed are estimates. Ranking does not establish task quality or production readiness.<\/p><dl><dt>Start with<\/dt><dd>I have hardware<\/dd><dt>Task<\/dt><dd>General<\/dd><dt>Application language<\/dt><dd>Any<\/dd><dt>Number of identical devices<\/dt><dd>1<\/dd><dt>Distribution<\/dt><dd>Single device<\/dd><dt>KV cache: bits per value<\/dt><dd>16<\/dd><dt>KV cache: location<\/dt><dd>GPU \/ unified memory<\/dd><dt>Additional memory (workspace), minimum per device in GiB<\/dt><dd>1<\/dd><dt>Additional memory (workspace), maximum per device in GiB<\/dt><dd>3<\/dd><dt>Host RAM (GiB)<\/dt><dd>64<\/dd><dt>Consider CPU offload<\/dt><dd>No<\/dd><dt>Process images \/ scanned documents<\/dt><dd>No<\/dd><dt>Require tool use<\/dt><dd>No<\/dd><dt>For business use<\/dt><dd>No<\/dd><dt>Independent tests for the selected language<\/dt><dd>No<\/dd><\/dl><p>Calculation: memory by memory-envelope-v1, ranking by shortlist-rank-v3.<\/p><section class=\"tk-llm-decision-candidate\"><h5>Devstral-Small-2-24B-Instruct-2512<\/h5><p>GeForce RTX 4090 \u00b7 Q4_K_M \u00b7 llama.cpp b10964<\/p><p>Fits into memory with headroom. Estimated 37 to 52 tokens per second at the selected context. Test the suitability with 20 questions of your own whose answer you know.<\/p><ul><li>Fits the memory estimate, 3.5 GiB stay free<\/li><li>Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li>Suitability: no independent data<\/li><\/ul><p>Required per device: 15.6 to 18.1 GiB \u00b7 Available per device: 21.6 GiB<\/p><p>Estimated generation: 37 to 52 tokens\/s. Measure time to first token and throughput under load separately.<\/p><p>Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.<\/p><\/section><section class=\"tk-llm-decision-candidate\"><h5>Magistral-Small-2509<\/h5><p>GeForce RTX 4090 \u00b7 Q4_K_M \u00b7 llama.cpp b10964<\/p><p>Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.<\/p><ul><li>Fits the memory estimate, 3.5 GiB stay free<\/li><li>Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li>Suitability: no independent data<\/li><\/ul><p>Required per device: 15.6 to 18.1 GiB \u00b7 Available per device: 21.6 GiB<\/p><p>Estimated generation: 37 to 52 tokens\/s. Measure time to first token and throughput under load separately.<\/p><p>Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.<\/p><\/section><section class=\"tk-llm-decision-candidate\"><h5>Mistral-Small-3.2-24B-Instruct-2506<\/h5><p>GeForce RTX 4090 \u00b7 Q4_K_M \u00b7 llama.cpp b10964<\/p><p>Comes directly after the first test candidate in the ranking. The place says nothing about answer quality.<\/p><ul><li>Fits the memory estimate, 3.5 GiB stay free<\/li><li>Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li>Suitability: vendor data only<\/li><\/ul><p>Required per device: 15.6 to 18.1 GiB \u00b7 Available per device: 21.6 GiB<\/p><p>Estimated generation: 37 to 52 tokens\/s. Measure time to first token and throughput under load separately.<\/p><p>Start the named configuration on the target device. Record model revision, runtime version, peak memory and execution errors.<\/p><\/section><h5>Acceptance with your own tasks<\/h5><p>For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.<\/p><p>Suggested pilot: 20 representative cases with expected answers defined in advance, including failure cases and questions with no supported answer. Use the same prompts and documents for every candidate. Set minimum quality and maximum waiting time before testing. Record task scores, citation errors, time to first token and total time per case. Repeat under planned concurrency.<\/p><p>The result link opens this calculation with its assumptions and sources. Catalogue updates may change later results. Also save the JSON export with sources and snapshot for an auditable record.<\/p><p><a class=\"tk-llm-share\" href=\"https:\/\/www.torck.io\/en\/llm-finder\/rtx-4090\/?llm%5Bmode%5D=hardware&amp;llm%5Btask%5D=general&amp;llm%5Blanguage%5D=any&amp;llm%5Bsort%5D=fit&amp;llm%5Brank%5D=v3&amp;llm%5Bparallel%5D=single&amp;llm%5Bkv_location%5D=gpu&amp;llm%5Bresult_filter%5D=all&amp;llm%5Bhardware%5D=rtx4090&amp;llm%5Bmodel%5D=&amp;llm%5Bruntime%5D=&amp;llm%5Bquantization%5D=&amp;llm%5Bcontext%5D=8192&amp;llm%5Bconcurrency%5D=1&amp;llm%5Bgpus%5D=1&amp;llm%5Bprefill_batch%5D=512&amp;llm%5Bkv_bits%5D=16&amp;llm%5Breserve_percent%5D=10&amp;llm%5Bworkspace_low%5D=1&amp;llm%5Bworkspace_high%5D=3&amp;llm%5Bhost_ram%5D=64&amp;llm%5Bcustom_memory_gib%5D=0&amp;llm%5Bpage%5D=1&amp;llm%5Bper_page%5D=20&amp;llm%5Boffload%5D=&amp;llm%5Bcommercial%5D=&amp;llm%5Bindependent_language%5D=&amp;llm%5Bvision%5D=&amp;llm%5Btools%5D=&amp;llm%5Brun%5D=1\">Result link to share<\/a><\/p><\/div><div class=\"tk-llm-decision-actions\"><button type=\"button\" data-decision-download hidden>Download brief (printable HTML)<\/button> <button type=\"button\" data-test-download hidden>Download test worksheet (CSV)<\/button><\/div><script type=\"application\/json\" class=\"tk-llm-test-template\">\"Case,Task,Expected answer,Model revision,Runtime and version,Task passed,Citation errors,First token ms,Total time ms,Notes\\r\\n1,,,,,,,,,\\r\\n2,,,,,,,,,\\r\\n3,,,,,,,,,\\r\\n4,,,,,,,,,\\r\\n5,,,,,,,,,\\r\\n6,,,,,,,,,\\r\\n7,,,,,,,,,\\r\\n8,,,,,,,,,\\r\\n9,,,,,,,,,\\r\\n10,,,,,,,,,\\r\\n11,,,,,,,,,\\r\\n12,,,,,,,,,\\r\\n13,,,,,,,,,\\r\\n14,,,,,,,,,\\r\\n15,,,,,,,,,\\r\\n16,,,,,,,,,\\r\\n17,,,,,,,,,\\r\\n18,,,,,,,,,\\r\\n19,,,,,,,,,\\r\\n20,,,,,,,,,\"<\/script><\/details><details class=\"tk-llm-cost\"><summary>Local deployment and API: your cost estimate<\/summary><p>All amounts are your own assumptions in EUR excluding tax. Fields remain on this device. Downloading the decision brief includes them in the file. Enter 0 for costs that do not apply. Compare only options meeting your quality, availability and data handling needs.<\/p><div class=\"tk-llm-grid\"><label>Purchase and setup (EUR)<input type=\"number\" data-cost=\"purchase\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>Useful life (months)<input type=\"number\" data-cost=\"months\" min=\"1\" step=\"any\" inputmode=\"decimal\"><\/label><label>Average system power (W)<input type=\"number\" data-cost=\"watts\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>Operating hours per month<input type=\"number\" data-cost=\"hours\" min=\"0\" max=\"744\" step=\"any\" inputmode=\"decimal\"><\/label><label>Electricity (EUR\/kWh)<input type=\"number\" data-cost=\"energy\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>Operations, staff and other costs (EUR\/month)<input type=\"number\" data-cost=\"operations\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>Requests per month<input type=\"number\" data-cost=\"requests\" min=\"1\" step=\"any\" inputmode=\"decimal\"><\/label><label>Input tokens per request<input type=\"number\" data-cost=\"input_tokens\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>Output tokens per request<input type=\"number\" data-cost=\"output_tokens\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>API input price (EUR per million tokens)<input type=\"number\" data-cost=\"input_price\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><label>API output price (EUR per million tokens)<input type=\"number\" data-cost=\"output_price\" min=\"0\" step=\"any\" inputmode=\"decimal\"><\/label><\/div><p>Local\/month = purchase \u00f7 useful life + W \u00f7 1000 \u00d7 hours \u00d7 electricity price + operations. API\/month = requests \u00d7 (input tokens \u00d7 input price + output tokens \u00d7 output price) \u00f7 1,000,000. Account separately for retrieval, storage, redundancy and API base fees. This estimate checks neither capacity nor comparable answer quality.<\/p><output class=\"tk-llm-cost-result\" aria-live=\"polite\" data-template=\"Local: {local} EUR\/month ({unit} EUR\/request). API: {api} EUR\/month.\" data-missing=\"Complete all fields with valid values to compare costs.\">Complete all fields with valid values to compare costs.<\/output><\/details><section class=\"tk-llm-result-notes\" aria-labelledby=\"tk-llm-notes-title\"><h4 id=\"tk-llm-notes-title\">Applies to all results<\/h4><p>Context: 8,192 tokens. Concurrent requests: 1. Reserve: 10%.<\/p><p>Speed is estimated. Only a test with your own examples shows whether a model handles your tasks well. Memory values are calculated, not measured. The license information does not replace a legal review for your intended use.<\/p><div class=\"tk-llm-quality-test\" id=\"tk-llm-quality-test\" tabindex=\"-1\"><p class=\"tk-llm-why-title\">Test suitability for your task<\/p><p>For the selection, what counts is whether the model answers your typical questions correctly. Ask it real questions from your business whose answers you already know. Then count the answers that are correct and complete.<\/p><\/div><\/section><details class=\"tk-llm-page-method\" id=\"tk-llm-method\"><summary>How the finder calculates<\/summary><p>Upper bound: memory bandwidth divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache at the assumed fill level. The range is 60 to 80 percent of this upper bound. This is a rule of thumb, not a guarantee of accuracy.<\/p><p>Bytes per weight come from the file size of the model where the catalogue has it, otherwise from the bit width of the quantization. The memory demand uses the same figure.<\/p><p>The range applies to one request, without processing the prompt. Which runtime is meant stands on each card.<\/p><ul><li>Memory bandwidth 1,010 GB\/s: <a href=\"https:\/\/zenodo.org\/records\/20390790\" target=\"_blank\" rel=\"noopener noreferrer\">Max Vyaznikov, GPU Ark (Zenodo, doi:10.5281\/zenodo.20390790)<\/a> \u00b7 Sep 17, 2026<\/li><li>Memory bandwidth 1,008 GB\/s: <a href=\"https:\/\/www.nvidia.com\/en-sg\/geforce\/graphics-cards\/50-series\/rtx-5090\/\" target=\"_blank\" rel=\"noopener noreferrer\">NVIDIA<\/a> \u00b7 Sep 17, 2026<\/li><\/ul><p>Prompt processing before the first answer is not estimated. Backend, driver and build can change the real rate considerably.<\/p><p class=\"tk-llm-small\">\u201cOfficially built\u201d says nothing about speed or correct output.<\/p><\/details><section class=\"tk-llm-list\" aria-labelledby=\"tk-llm-list-title\"><h4 id=\"tk-llm-list-title\">All 1,092 configurations<\/h4>\n<ul class=\"tk-llm-cards\" role=\"list\"><li id=\"tk-llm-row-f29fc6161377\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-f29fc6161377-title\"><h5 id=\"tk-llm-row-f29fc6161377-title\">Devstral-Small-2-24B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Nov 28, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Devstral-Small-2-24B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Devstral-Small-2-24B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Devstral-Small-2-24B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 24.0 billion parameters, 262,144 tokens of context length and the creation date Nov 28, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Devstral-Small-2-24B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Devstral-Small-2-24B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-f29fc6161377\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-f29fc6161377\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-f1609d8fb204\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-f1609d8fb204-title\"><h5 id=\"tk-llm-row-f1609d8fb204-title\">Magistral-Small-2509<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Sep 12, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2509\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Magistral-Small-2509<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Magistral-Small-2509 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Sep 12, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2509\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Magistral-Small-2509 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-f1609d8fb204\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-f1609d8fb204\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-50468b18e01e\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-50468b18e01e-title\"><h5 id=\"tk-llm-row-50468b18e01e-title\">Mistral-Small-3.2-24B-Instruct-2506<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Jun 19, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.2-24B-Instruct-2506\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Mistral-Small-3.2-24B-Instruct-2506<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-info\"><span aria-hidden=\"true\">i<\/span> Suitability: vendor data only<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><div class=\"tk-llm-evaluation\"><span class=\"tk-llm-quality-state\">Vendor claim<\/span><p>MMLU Pro: 69.1%<\/p><small>Figure for the model variant. No measurement is available for this quantization.<\/small><small>Vendor-reported with five examples and chain of thought. Not independently reproduced.<\/small><small><a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.2-24B-Instruct-2506\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face Hub \u00b7 Jun 20, 2025<\/a><\/small><\/div><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-50468b18e01e\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-50468b18e01e\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-b5085f58974e\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-b5085f58974e-title\"><h5 id=\"tk-llm-row-b5085f58974e-title\">Magistral-Small-2506<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Jun 4, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2506\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Magistral-Small-2506<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Magistral-Small-2506 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 23.6 billion parameters, 40,960 tokens of context length and the creation date Jun 4, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Magistral-Small-2506\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Magistral-Small-2506 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-b5085f58974e\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-b5085f58974e\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-cfca3f9abd56\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-cfca3f9abd56-title\"><h5 id=\"tk-llm-row-cfca3f9abd56-title\">Mistral-Small-3.1-24B-Instruct-2503<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Mar 11, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.1-24B-Instruct-2503\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Mistral-Small-3.1-24B-Instruct-2503<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 57 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Mistral-Small-3.1-24B-Instruct-2503 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 24.0 billion parameters, 131,072 tokens of context length and the creation date Mar 11, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Mistral-Small-3.1-24B-Instruct-2503\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Mistral-Small-3.1-24B-Instruct-2503 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-cfca3f9abd56\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-cfca3f9abd56\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-f4cda5073182\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-f4cda5073182-title\"><h5 id=\"tk-llm-row-f4cda5073182-title\">Codestral-22B-v0.1<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on May 29, 2024 <a href=\"https:\/\/huggingface.co\/mistralai\/Codestral-22B-v0.1\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Codestral-22B-v0.1<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 4.0 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 38 to 54 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.2 to 17.6 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 43 to 61 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Codestral-22B-v0.1 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 22.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2024. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Codestral-22B-v0.1\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Codestral-22B-v0.1 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-f4cda5073182\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-f4cda5073182\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li class=\"tk-llm-stream-cta\"><div class=\"tk-inline-cta tk-llm-stream-contact\"><p><b>Test with your own tasks<\/b>The list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.<\/p><p class=\"tk-llm-stream-action\"><a class=\"tk-llm-contact-link\" href=\"\/en\/#contact-us\" aria-haspopup=\"dialog\" data-tk-action=\"contact-stream\">Discuss a test<\/a><\/p><\/div><\/li><li id=\"tk-llm-row-08be609e4710\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-08be609e4710-title\"><h5 id=\"tk-llm-row-08be609e4710-title\">Ministral-3-14B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-14B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 9.4 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 61 to 85 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>9.9 to 12.2 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 70 to 98 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-14B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-08be609e4710\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-08be609e4710\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-ce4167791baf\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-ce4167791baf-title\"><h5 id=\"tk-llm-row-ce4167791baf-title\">Ministral-3-14B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-14B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 56 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-14B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-14B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-ce4167791baf\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-ce4167791baf\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-7be21c3b91b7\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-7be21c3b91b7-title\"><h5 id=\"tk-llm-row-7be21c3b91b7-title\">Ministral-3-14B-Reasoning-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-14B-Reasoning-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 9.4 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 61 to 85 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>9.9 to 12.2 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 70 to 98 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-7be21c3b91b7\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-7be21c3b91b7\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-92ba3dfd213e\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-92ba3dfd213e-title\"><h5 id=\"tk-llm-row-92ba3dfd213e-title\">Ministral-3-14B-Reasoning-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-14B-Reasoning-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 3.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 37 to 52 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>15.6 to 18.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 40 to 56 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-14B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 13.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-14B-Reasoning-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-92ba3dfd213e\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-92ba3dfd213e\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-839e6359f2b0\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-839e6359f2b0-title\"><h5 id=\"tk-llm-row-839e6359f2b0-title\">Ministral-3-8B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-8B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 12.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 93 to 128 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>6.9 to 9.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 111 to 154 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-8B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-839e6359f2b0\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-839e6359f2b0\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-6a888d06a390\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-6a888d06a390-title\"><h5 id=\"tk-llm-row-6a888d06a390-title\">Ministral-3-8B-Instruct-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-8B-Instruct-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 8.8 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 57 to 80 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>10.5 to 12.8 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 64 to 89 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-8B-Instruct-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-8B-Instruct-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-6a888d06a390\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-6a888d06a390\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li class=\"tk-llm-stream-cta\"><div class=\"tk-inline-cta tk-llm-stream-contact\"><p><b>Test with your own tasks<\/b>The list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.<\/p><p class=\"tk-llm-stream-action\"><a class=\"tk-llm-contact-link\" href=\"\/en\/#contact-us\" aria-haspopup=\"dialog\" data-tk-action=\"contact-stream\">Discuss a test<\/a><\/p><\/div><\/li><li id=\"tk-llm-row-91e10ed8aca3\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-91e10ed8aca3-title\"><h5 id=\"tk-llm-row-91e10ed8aca3-title\">Ministral-3-8B-Reasoning-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-8B-Reasoning-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 12.5 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 93 to 128 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>6.9 to 9.1 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 111 to 154 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-91e10ed8aca3\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-91e10ed8aca3\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-5c9f8eb453d7\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-5c9f8eb453d7-title\"><h5 id=\"tk-llm-row-5c9f8eb453d7-title\">Ministral-3-8B-Reasoning-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-8B-Reasoning-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 8.8 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 57 to 80 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>10.5 to 12.8 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 64 to 89 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-5c9f8eb453d7\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-5c9f8eb453d7\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-827ccbaaa24c\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-827ccbaaa24c-title\"><h5 id=\"tk-llm-row-827ccbaaa24c-title\">Ministral-3-8B-Reasoning-2512<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>BF16<\/dd><\/div><div><dt>Runtime<\/dt><dd>vLLM<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Oct 31, 2025 <a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Ministral-3-8B-Reasoning-2512<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-warn\"><span aria-hidden=\"true\">!<\/span> Fits with little headroom, 0.4 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 30 to 43 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-tight\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>18.7 to 21.2 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 32 to 46 tokens per second with a short context (512 tokens).<\/p><small>The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.<\/small><small>With vLLM, under load or with multi-token prediction, real values can differ considerably.<\/small><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For Ministral-3-8B-Reasoning-2512 the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.9 billion parameters, 16,384 tokens of context length and the creation date Oct 31, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Reasoning-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of Ministral-3-8B-Reasoning-2512 at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-827ccbaaa24c\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-827ccbaaa24c\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-7e8fe8cc0731\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-7e8fe8cc0731-title\"><h5 id=\"tk-llm-row-7e8fe8cc0731-title\">DeepSeek-R1-0528-Qwen3-8B<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on May 29, 2025 <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ DeepSeek-R1-0528-Qwen3-8B<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 12.6 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 94 to 130 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>6.8 to 9.0 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 115 to 159 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-7e8fe8cc0731\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-7e8fe8cc0731\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-4dc3881c5a93\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-4dc3881c5a93-title\"><h5 id=\"tk-llm-row-4dc3881c5a93-title\">DeepSeek-R1-0528-Qwen3-8B<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on May 29, 2025 <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ DeepSeek-R1-0528-Qwen3-8B<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 9.1 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 59 to 82 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>10.2 to 12.5 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 66 to 92 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-4dc3881c5a93\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-4dc3881c5a93\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-837aea297dca\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-837aea297dca-title\"><h5 id=\"tk-llm-row-837aea297dca-title\">DeepSeek-R1-0528-Qwen3-8B<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>BF16<\/dd><\/div><div><dt>Runtime<\/dt><dd>vLLM<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on May 29, 2025 <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ DeepSeek-R1-0528-Qwen3-8B<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-warn\"><span aria-hidden=\"true\">!<\/span> Fits with little headroom, 1.7 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 33 to 46 tokens per second at 8,192 tokens of context<\/li><li class=\"is-open\"><span aria-hidden=\"true\">?<\/span> Suitability: no independent data<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-tight\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>17.4 to 19.9 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 35 to 50 tokens per second with a short context (512 tokens).<\/p><small>The efficiency corridor of the estimate comes from a measurement series on llama.cpp. It is not verified for other runtimes.<\/small><small>With vLLM, under load or with multi-token prediction, real values can differ considerably.<\/small><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><p>For DeepSeek-R1-0528-Qwen3-8B the catalogue holds no publisher measurement for the task General.<\/p><p>The catalogue names 8.2 billion parameters, 32,768 tokens of context length and the creation date May 29, 2025. These figures bound the fit, they do not prove it.<\/p><p><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-0528-Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">Open the model card of DeepSeek-R1-0528-Qwen3-8B at the publisher<\/a><\/p><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-837aea297dca\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-837aea297dca\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li class=\"tk-llm-stream-cta\"><div class=\"tk-inline-cta tk-llm-stream-contact\"><p><b>Test with your own tasks<\/b>The list shows calculated values. Whether a model solves your tasks only shows in a test with your own data.<\/p><p class=\"tk-llm-stream-action\"><a class=\"tk-llm-contact-link\" href=\"\/en\/#contact-us\" aria-haspopup=\"dialog\" data-tk-action=\"contact-stream\">Discuss a test<\/a><\/p><\/div><\/li><li id=\"tk-llm-row-9b2af98e677c\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-9b2af98e677c-title\"><h5 id=\"tk-llm-row-9b2af98e677c-title\">Qwen3-8B<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q4_K_M<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Apr 27, 2025 <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Qwen3-8B<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 12.6 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 94 to 130 tokens per second at 8,192 tokens of context<\/li><li class=\"is-info\"><span aria-hidden=\"true\">i<\/span> Suitability: vendor data only<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>6.8 to 9.0 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 115 to 159 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><div class=\"tk-llm-evaluation\"><span class=\"tk-llm-quality-state\">Vendor claim<\/span><p>IFEval strict prompt \/ non-thinking: 83.0%<\/p><small>Figure for the model variant. No measurement is available for this quantization.<\/small><small>Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.<\/small><small><a href=\"https:\/\/arxiv.org\/html\/2505.09388\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) \u00b7 May 14, 2025<\/a><\/small><\/div><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-9b2af98e677c\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-9b2af98e677c\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><li id=\"tk-llm-row-3f4a8501d8b1\"><article class=\"tk-llm-card\" aria-labelledby=\"tk-llm-row-3f4a8501d8b1-title\"><h5 id=\"tk-llm-row-3f4a8501d8b1-title\">Qwen3-8B<\/h5><dl class=\"tk-llm-card-meta\"><div><dt>Quantization<\/dt><dd>Q8_0<\/dd><\/div><div><dt>Runtime<\/dt><dd>llama.cpp<\/dd><\/div><\/dl><p class=\"tk-llm-released\">Repository created on Apr 27, 2025 <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-8B\" target=\"_blank\" rel=\"noopener noreferrer\">huggingface.co \/ Qwen3-8B<\/a><\/p><ul class=\"tk-llm-verdict\" role=\"list\" aria-label=\"Summary\"><li class=\"is-ok\"><span aria-hidden=\"true\">\u2713<\/span> Fits the memory estimate, 9.1 GiB stay free<\/li><li class=\"is-estimate\"><span aria-hidden=\"true\">\u2248<\/span> Estimated 59 to 82 tokens per second at 8,192 tokens of context<\/li><li class=\"is-info\"><span aria-hidden=\"true\">i<\/span> Suitability: vendor data only<\/li><\/ul><div class=\"tk-llm-dims\"><div class=\"tk-llm-dim\"><h6>Memory fit<\/h6><div class=\"tk-llm-dim-tone is-fits\"><dl class=\"tk-llm-figures\"><dt>Required per device<\/dt><dd>10.2 to 12.5 GiB<\/dd><dt>Available per device<\/dt><dd>21.6 GiB<\/dd><\/dl><\/div><\/div><div class=\"tk-llm-dim\"><h6>Response speed<\/h6><p class=\"tk-llm-speed-range\">About 66 to 92 tokens per second with a short context (512 tokens).<\/p><\/div><div class=\"tk-llm-dim\"><h6>Suitability for your task<\/h6><div class=\"tk-llm-evaluation\"><span class=\"tk-llm-quality-state\">Vendor claim<\/span><p>IFEval strict prompt \/ non-thinking: 83.0%<\/p><small>Figure for the model variant. No measurement is available for this quantization.<\/small><small>Non-thinking mode. Strict-prompt instruction accuracy. Manufacturer result from May 2025.<\/small><small><a href=\"https:\/\/arxiv.org\/html\/2505.09388\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen Team, Qwen3 Technical Report (arXiv:2505.09388) \u00b7 May 14, 2025<\/a><\/small><\/div><p><a href=\"#tk-llm-quality-test\">Test suitability yourself<\/a><\/p><\/div><\/div><details class=\"tk-llm-why\" data-evidence=\"tk-llm-row-3f4a8501d8b1\" data-evidence-page=\"1\"><summary>Why this assessment<\/summary><p class=\"tk-llm-why-load\"><a href=\"?llm%5Bevidence%5D=1&amp;llm%5Brun%5D=1#tk-llm-row-3f4a8501d8b1\">Show the calculation and its sources<\/a><\/p><\/details><\/article><\/li><\/ul><\/section>\n<div class=\"tk-llm-load\" hidden><button type=\"button\" class=\"tk-llm-load-more\" data-load-page=\"2\" data-load-per-page=\"20\">Load more<\/button><p class=\"tk-llm-load-status\" data-load-shown=\"20\" data-load-total=\"1092\">20 of 1,092 shown<\/p><\/div><nav class=\"tk-llm-pagination\" aria-label=\"Result pages\"><a class=\"tk-llm-page\" data-page=\"2\" href=\"?llm%5Bmode%5D=hardware&amp;llm%5Btask%5D=general&amp;llm%5Blanguage%5D=any&amp;llm%5Bsort%5D=fit&amp;llm%5Brank%5D=v3&amp;llm%5Bparallel%5D=single&amp;llm%5Bkv_location%5D=gpu&amp;llm%5Bresult_filter%5D=all&amp;llm%5Bhardware%5D=rtx4090&amp;llm%5Bmodel%5D=&amp;llm%5Bruntime%5D=&amp;llm%5Bquantization%5D=&amp;llm%5Bcontext%5D=8192&amp;llm%5Bconcurrency%5D=1&amp;llm%5Bgpus%5D=1&amp;llm%5Bprefill_batch%5D=512&amp;llm%5Bkv_bits%5D=16&amp;llm%5Breserve_percent%5D=10&amp;llm%5Bworkspace_low%5D=1&amp;llm%5Bworkspace_high%5D=3&amp;llm%5Bhost_ram%5D=64&amp;llm%5Bcustom_memory_gib%5D=0&amp;llm%5Bpage%5D=2&amp;llm%5Bper_page%5D=20&amp;llm%5Boffload%5D=&amp;llm%5Bcommercial%5D=&amp;llm%5Bindependent_language%5D=&amp;llm%5Bvision%5D=&amp;llm%5Btools%5D=&amp;llm%5Brun%5D=1\">Next<\/a><span>Page 1 of 55<\/span><\/nav><button type=\"button\" class=\"tk-llm-export\" hidden>Save results as JSON<\/button>\n<script type=\"application\/json\" class=\"tk-llm-contact-draft\">\"LLM Finder: test with our own tasks\\n\\nHello, we would like to test the following models with our own tasks.\\n\\nApproach: I have hardware\\nDevice: GeForce RTX 4090\\nContext: 8,192 tokens, concurrent requests: 1\\nData as of: Sep 19, 2026\\n\\nModels for the first test:\\n- Devstral-Small-2-24B-Instruct-2512, Q4_K_M, llama.cpp, GeForce RTX 4090: Fits the memory estimate, 3.5 GiB stay free. Estimated 37 to 52 tokens per second at 8,192 tokens of context. Suitability: no independent data.\\n- Magistral-Small-2509, Q4_K_M, llama.cpp, GeForce RTX 4090: Fits the memory estimate, 3.5 GiB stay free. Estimated 37 to 52 tokens per second at 8,192 tokens of context. Suitability: no independent data.\\n- Mistral-Small-3.2-24B-Instruct-2506, Q4_K_M, llama.cpp, GeForce RTX 4090: Fits the memory estimate, 3.5 GiB stay free. Estimated 37 to 52 tokens per second at 8,192 tokens of context. Suitability: vendor data only.\"<\/script><aside class=\"tk-inline-cta tk-llm-contact\" aria-label=\"Test with your own tasks\"><p><b>Test with your own tasks<\/b>The estimate above does not replace a test with your own data. torck runs this test with you.<\/p><div class=\"tk-llm-contact-actions\"><a class=\"tk-btn\" href=\"\/en\/#contact-us\" aria-haspopup=\"dialog\" data-tk-action=\"contact-block\">Discuss a test<\/a><p class=\"tk-llm-contact-trust\">An answer within one business day, straight from a developer.<\/p><\/div><\/aside><\/div>\n <section class=\"tk-llm-landing\" translate=\"no\" lang=\"en\"><h2>How to check the result on an RTX 4090<\/h2><p>Model weights, quantization and memory for active requests determine which configurations fit an RTX 4090. The calculator uses catalogue hardware data with sources in each configuration\u2019s details.<\/p><p>This example uses 8,192 context tokens and one concurrent request. It includes a ten percent memory reserve and the configured runtime workspace. Increasing context or concurrency can change the selection.<\/p><p>Check the upper memory bound first. Then compare estimated generation speed and measure your model file on the target device. Time to first token and throughput with multiple users require separate measurements.<\/p><p>A memory calculation alone is insufficient for procurement. Use the decision brief to test the same business tasks with each candidate. Check licences, runtime support and outstanding evidence per configuration.<\/p><\/section>  <details class=\"tk-llm-catalog\" id=\"tk-llm-catalog\"><summary>Model information<\/summary><p>The catalog lists parameters, context length, license, and source. A catalog entry does not confirm that a model will run on your hardware.<\/p><p>The catalogue holds 1,081 models with technical data, 3,958 more by name only and 1,154 format editions of those models. The search finds all three.<\/p><form class=\"tk-llm-model-search\" method=\"get\" action=\"https:\/\/www.torck.io\/en\/llm-finder\/rtx-4090\/\"><input type=\"hidden\" name=\"llm[mode]\" value=\"hardware\"><input type=\"hidden\" name=\"llm[task]\" value=\"general\"><input type=\"hidden\" name=\"llm[language]\" value=\"any\"><input type=\"hidden\" name=\"llm[sort]\" value=\"fit\"><input type=\"hidden\" name=\"llm[rank]\" value=\"v3\"><input type=\"hidden\" name=\"llm[parallel]\" value=\"single\"><input type=\"hidden\" name=\"llm[kv_location]\" value=\"gpu\"><input type=\"hidden\" name=\"llm[result_filter]\" value=\"all\"><input type=\"hidden\" name=\"llm[hardware]\" value=\"rtx4090\"><input type=\"hidden\" name=\"llm[model]\" value=\"\"><input type=\"hidden\" name=\"llm[runtime]\" value=\"\"><input type=\"hidden\" name=\"llm[quantization]\" value=\"\"><input type=\"hidden\" name=\"llm[context]\" value=\"8192\"><input type=\"hidden\" name=\"llm[concurrency]\" value=\"1\"><input type=\"hidden\" name=\"llm[gpus]\" value=\"1\"><input type=\"hidden\" name=\"llm[prefill_batch]\" value=\"512\"><input type=\"hidden\" name=\"llm[kv_bits]\" value=\"16\"><input type=\"hidden\" name=\"llm[reserve_percent]\" value=\"10\"><input type=\"hidden\" name=\"llm[workspace_low]\" value=\"1\"><input type=\"hidden\" name=\"llm[workspace_high]\" value=\"3\"><input type=\"hidden\" name=\"llm[host_ram]\" value=\"64\"><input type=\"hidden\" name=\"llm[custom_memory_gib]\" value=\"0\"><input type=\"hidden\" name=\"llm[page]\" value=\"1\"><input type=\"hidden\" name=\"llm[per_page]\" value=\"20\"><input type=\"hidden\" name=\"llm[offload]\" value=\"\"><input type=\"hidden\" name=\"llm[commercial]\" value=\"\"><input type=\"hidden\" name=\"llm[independent_language]\" value=\"\"><input type=\"hidden\" name=\"llm[vision]\" value=\"\"><input type=\"hidden\" name=\"llm[tools]\" value=\"\"><input type=\"hidden\" name=\"llm[run]\" value=\"1\"><label for=\"tk-llm-field-q\">Search the catalogue for a model<input type=\"search\" id=\"tk-llm-field-q\" name=\"llm[q]\" value=\"\" maxlength=\"120\" autocomplete=\"off\" aria-describedby=\"tk-llm-help-q\"><\/label><small id=\"tk-llm-help-q\">Enter part of a name, for example Mellum, xLAM or Qwen3-Omni. Capitals and hyphens do not matter.<\/small><button type=\"submit\">Search<\/button><\/form><div class=\"tk-llm-model-search-out\" role=\"region\" aria-live=\"polite\"><\/div><h3 class=\"tk-llm-directory-title\">Publisher directory<\/h3><p>The finder reads the publishers\u2019 model accounts in full. This overview shows per account how many models it computes with data and how many it only lists so far. The individual names are in the search above.<\/p><p>1,081 models with data, 3,958 more listed only, from 53 publisher accounts.<\/p><p>For the models that are only listed, the next reconciliation run fetches the technical data. Until then you find the name through the search and the model through the link to the account.<\/p><p>On top of that there are 1,154 format editions. That is the same model as GGUF, MLX, FP8, NVFP4, AWQ or GPTQ. The finder calculates them as a configuration of the original and lists their name so the search finds it.<\/p><div class=\"tk-llm-table\" role=\"region\" tabindex=\"0\" aria-label=\"Publisher directory\"><table><thead><tr><th scope=\"col\">Publisher<\/th><th scope=\"col\">With data<\/th><th scope=\"col\">Listed only<\/th><th scope=\"col\">Format editions<\/th><\/tr><\/thead><tbody><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/01-ai\" target=\"_blank\" rel=\"noopener noreferrer\">01-ai<\/a><\/th><td>6<\/td><td>22<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/ai21labs\" target=\"_blank\" rel=\"noopener noreferrer\">ai21labs<\/a><\/th><td>11<\/td><td>1<\/td><td>4<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/Aleph-Alpha\" target=\"_blank\" rel=\"noopener noreferrer\">Aleph-Alpha<\/a><\/th><td>7<\/td><td>13<\/td><td>3<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/allenai\" target=\"_blank\" rel=\"noopener noreferrer\">allenai<\/a><\/th><td>25<\/td><td>792<\/td><td>20<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/apple\" target=\"_blank\" rel=\"noopener noreferrer\">apple<\/a><\/th><td>5<\/td><td>49<\/td><td>2<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/arcee-ai\" target=\"_blank\" rel=\"noopener noreferrer\">arcee-ai<\/a><\/th><td>20<\/td><td>101<\/td><td>59<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/baichuan-inc\" target=\"_blank\" rel=\"noopener noreferrer\">baichuan-inc<\/a><\/th><td>5<\/td><td>8<\/td><td>5<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/baidu\" target=\"_blank\" rel=\"noopener noreferrer\">baidu<\/a><\/th><td>14<\/td><td>3<\/td><td>1<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/ByteDance-Seed\" target=\"_blank\" rel=\"noopener noreferrer\">ByteDance-Seed<\/a><\/th><td>12<\/td><td>28<\/td><td>2<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/CohereLabs\" target=\"_blank\" rel=\"noopener noreferrer\">CohereLabs<\/a><\/th><td>16<\/td><td>14<\/td><td>12<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/deepseek-ai\" target=\"_blank\" rel=\"noopener noreferrer\">deepseek-ai<\/a><\/th><td>68<\/td><td>48<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/google\" target=\"_blank\" rel=\"noopener noreferrer\">google<\/a><\/th><td>66<\/td><td>749<\/td><td>37<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/HuggingFaceTB\" target=\"_blank\" rel=\"noopener noreferrer\">HuggingFaceTB<\/a><\/th><td>22<\/td><td>34<\/td><td>10<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/ibm-granite\" target=\"_blank\" rel=\"noopener noreferrer\">ibm-granite<\/a><\/th><td>56<\/td><td>75<\/td><td>60<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/IFM\" target=\"_blank\" rel=\"noopener noreferrer\">IFM<\/a><\/th><td>12<\/td><td>14<\/td><td>9<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/inclusionAI\" target=\"_blank\" rel=\"noopener noreferrer\">inclusionAI<\/a><\/th><td>23<\/td><td>118<\/td><td>37<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/internlm\" target=\"_blank\" rel=\"noopener noreferrer\">internlm<\/a><\/th><td>9<\/td><td>69<\/td><td>29<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/JetBrains\" target=\"_blank\" rel=\"noopener noreferrer\">JetBrains<\/a><\/th><td>3<\/td><td>11<\/td><td>16<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/kakaocorp\" target=\"_blank\" rel=\"noopener noreferrer\">kakaocorp<\/a><\/th><td>14<\/td><td>2<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/LGAI-EXAONE\" target=\"_blank\" rel=\"noopener noreferrer\">LGAI-EXAONE<\/a><\/th><td>13<\/td><td>9<\/td><td>28<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/LiquidAI\" target=\"_blank\" rel=\"noopener noreferrer\">LiquidAI<\/a><\/th><td>16<\/td><td>23<\/td><td>106<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/llm-jp\" target=\"_blank\" rel=\"noopener noreferrer\">llm-jp<\/a><\/th><td>7<\/td><td>242<\/td><td>3<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/meta-llama\" target=\"_blank\" rel=\"noopener noreferrer\">meta-llama<\/a><\/th><td>62<\/td><td>30<\/td><td>7<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/microsoft\" target=\"_blank\" rel=\"noopener noreferrer\">microsoft<\/a><\/th><td>83<\/td><td>234<\/td><td>29<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/MiniMaxAI\" target=\"_blank\" rel=\"noopener noreferrer\">MiniMaxAI<\/a><\/th><td>16<\/td><td>7<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/mistralai\" target=\"_blank\" rel=\"noopener noreferrer\">mistralai<\/a><\/th><td>33<\/td><td>17<\/td><td>16<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/moonshotai\" target=\"_blank\" rel=\"noopener noreferrer\">moonshotai<\/a><\/th><td>16<\/td><td>0<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/naver-hyperclovax\" target=\"_blank\" rel=\"noopener noreferrer\">naver-hyperclovax<\/a><\/th><td>4<\/td><td>2<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/NousResearch\" target=\"_blank\" rel=\"noopener noreferrer\">NousResearch<\/a><\/th><td>17<\/td><td>75<\/td><td>33<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/nvidia\" target=\"_blank\" rel=\"noopener noreferrer\">nvidia<\/a><\/th><td>29<\/td><td>423<\/td><td>111<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/occiglot\" target=\"_blank\" rel=\"noopener noreferrer\">occiglot<\/a><\/th><td>10<\/td><td>0<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/openai\" target=\"_blank\" rel=\"noopener noreferrer\">openai<\/a><\/th><td>5<\/td><td>13<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/openbmb\" target=\"_blank\" rel=\"noopener noreferrer\">openbmb<\/a><\/th><td>22<\/td><td>68<\/td><td>61<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/openGPT-X\" target=\"_blank\" rel=\"noopener noreferrer\">openGPT-X<\/a><\/th><td>3<\/td><td>0<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/OpenGVLab\" target=\"_blank\" rel=\"noopener noreferrer\">OpenGVLab<\/a><\/th><td>4<\/td><td>181<\/td><td>23<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/PleIAs\" target=\"_blank\" rel=\"noopener noreferrer\">PleIAs<\/a><\/th><td>12<\/td><td>7<\/td><td>5<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/Qwen\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen<\/a><\/th><td>86<\/td><td>111<\/td><td>239<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/rinna\" target=\"_blank\" rel=\"noopener noreferrer\">rinna<\/a><\/th><td>4<\/td><td>26<\/td><td>27<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/Salesforce\" target=\"_blank\" rel=\"noopener noreferrer\">Salesforce<\/a><\/th><td>10<\/td><td>104<\/td><td>5<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/sarvamai\" target=\"_blank\" rel=\"noopener noreferrer\">sarvamai<\/a><\/th><td>5<\/td><td>2<\/td><td>6<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/ServiceNow-AI\" target=\"_blank\" rel=\"noopener noreferrer\">ServiceNow-AI<\/a><\/th><td>7<\/td><td>1<\/td><td>1<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/Skywork\" target=\"_blank\" rel=\"noopener noreferrer\">Skywork<\/a><\/th><td>13<\/td><td>20<\/td><td>5<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/stabilityai\" target=\"_blank\" rel=\"noopener noreferrer\">stabilityai<\/a><\/th><td>16<\/td><td>21<\/td><td>3<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/stepfun-ai\" target=\"_blank\" rel=\"noopener noreferrer\">stepfun-ai<\/a><\/th><td>14<\/td><td>10<\/td><td>8<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/swiss-ai\" target=\"_blank\" rel=\"noopener noreferrer\">swiss-ai<\/a><\/th><td>13<\/td><td>0<\/td><td>9<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/tencent\" target=\"_blank\" rel=\"noopener noreferrer\">tencent<\/a><\/th><td>30<\/td><td>33<\/td><td>38<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/tiiuae\" target=\"_blank\" rel=\"noopener noreferrer\">tiiuae<\/a><\/th><td>27<\/td><td>26<\/td><td>59<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/trillionlabs\" target=\"_blank\" rel=\"noopener noreferrer\">trillionlabs<\/a><\/th><td>8<\/td><td>17<\/td><td>2<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/upstage\" target=\"_blank\" rel=\"noopener noreferrer\">upstage<\/a><\/th><td>16<\/td><td>7<\/td><td>1<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/utter-project\" target=\"_blank\" rel=\"noopener noreferrer\">utter-project<\/a><\/th><td>20<\/td><td>16<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/xai-org\" target=\"_blank\" rel=\"noopener noreferrer\">xai-org<\/a><\/th><td>2<\/td><td>0<\/td><td>0<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\" target=\"_blank\" rel=\"noopener noreferrer\">XiaomiMiMo<\/a><\/th><td>9<\/td><td>6<\/td><td>1<\/td><\/tr><tr><th scope=\"row\"><a href=\"https:\/\/huggingface.co\/zai-org\" target=\"_blank\" rel=\"noopener noreferrer\">zai-org<\/a><\/th><td>55<\/td><td>76<\/td><td>22<\/td><\/tr><\/tbody><\/table><\/div><\/details> <section class=\"tk-llm-dataset\" id=\"tk-llm-dataset\" aria-labelledby=\"tk-llm-dataset-title\"><h3 id=\"tk-llm-dataset-title\">Open catalogue for retrieval<\/h3><p>The data this page calculates from is available as an open interface. Every observation names its source, that source&#039;s license and the date of capture. Responses are JSON and paginated, no account is needed.<\/p><ul class=\"tk-llm-dataset-routes\" role=\"list\"><li><a href=\"\/wp-json\/torck-llm\/v1\/catalog\">Header with snapshot, counts per collection and terms of use<\/a><small>\/wp-json\/torck-llm\/v1\/catalog<\/small><\/li><li><a href=\"\/wp-json\/torck-llm\/v1\/catalog\/entities?kind=ModelVariant\">Entries of one collection with their current values<\/a><small>\/wp-json\/torck-llm\/v1\/catalog\/entities?kind=ModelVariant<\/small><\/li><li><a href=\"\/wp-json\/torck-llm\/v1\/catalog\/values\">Single values with origin, source, date and unit<\/a><small>\/wp-json\/torck-llm\/v1\/catalog\/values<\/small><\/li><li><a href=\"\/wp-json\/torck-llm\/v1\/catalog\/sources\">Sources with license, status and terms of use<\/a><small>\/wp-json\/torck-llm\/v1\/catalog\/sources<\/small><\/li><li><a href=\"\/wp-json\/torck-llm\/v1\/catalog\/changes?since=0\">Changes since an earlier snapshot<\/a><small>\/wp-json\/torck-llm\/v1\/catalog\/changes?since=0<\/small><\/li><\/ul><p class=\"tk-llm-small\">Only facts whose source permits sharing with attribution are served. The terms of the source also apply to any reuse; the exclusion rule is stated in the header. Data snapshot <time datetime=\"2026-09-19\">Sep 19, 2026<\/time>.<\/p><\/section> <details class=\"tk-llm-attribution\"><summary>Sources and licenses<\/summary><p>The finder takes single data points from these sources and turns them into memory and suitability assessments. The sources have not reviewed this evaluation.<\/p><ul><li><a href=\"https:\/\/www.amd.com\/en\/products\/graphics\/desktops\/radeon\/7000-series\/amd-radeon-rx-7900xtx.html\" target=\"_blank\" rel=\"noopener noreferrer\">AMD<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 20, 2026 \u00b7 Source documents: 15<\/small><\/li><li><a href=\"https:\/\/support.apple.com\/en-us\/111894\" target=\"_blank\" rel=\"noopener noreferrer\">Apple<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 21<\/small><\/li><li><a href=\"https:\/\/www.heise.de\/news\/Was-ich-ueber-lokale-KI-gelernt-habe-11433840.html\" target=\"_blank\" rel=\"noopener noreferrer\">heise online, Jan-Keno Janssen<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste 01-ai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste ai21labs<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste Aleph-Alpha<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste allenai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste apple<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste arcee-ai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste baichuan-inc<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste baidu<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste ByteDance-Seed<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste CohereLabs<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste deepseek-ai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste google<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste HuggingFaceTB<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste ibm-granite<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste IFM<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste inclusionAI<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste internlm<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste JetBrains<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste kakaocorp<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste LGAI-EXAONE<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste LiquidAI<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste llm-jp<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste meta-llama<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste microsoft<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste MiniMaxAI<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste mistralai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste moonshotai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste naver-hyperclovax<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste NousResearch<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste nvidia<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste occiglot<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste openai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste openbmb<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste openGPT-X<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste OpenGVLab<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste PleIAs<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste Qwen<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste rinna<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste Salesforce<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste sarvamai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste ServiceNow-AI<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste Skywork<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste stabilityai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste stepfun-ai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste swiss-ai<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste tencent<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste tiiuae<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste trillionlabs<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste upstage<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste utter-project<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste xai-org<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste XiaomiMiMo<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face, Modellliste zai-org<\/a><br><small>License: Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/huggingface.co\/terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">Hugging Face Hub<\/a><br><small>License: Apache 2.0; Gemma Terms of Use; Llama 3 Community License; Llama 3.1 Community License; Llama 3.2 Community License; Llama 3.3 Community License; MIT; MIT with the terms of the Llama base model; Publisher&#039;s own license (text at the source) \u00b7 Used under license and terms of use \u00b7 Retrieved Oct 1, 2026 \u00b7 Source documents: 3971<\/small><\/li><li><a href=\"https:\/\/www.intel.com\/content\/www\/us\/en\/products\/sku\/227955\/intel-arc-a770-graphics-8gb\/specifications.html\" target=\"_blank\" rel=\"noopener noreferrer\">Intel<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 7<\/small><\/li><li><a href=\"https:\/\/arxiv.org\/html\/2601.09527v1\" target=\"_blank\" rel=\"noopener noreferrer\">Knoop, Holtmann (arXiv:2601.09527)<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/pypi.org\/project\/llama-roofline\/\" target=\"_blank\" rel=\"noopener noreferrer\">llama-roofline, Manu Nicholas Jacob<\/a><br><small>License: MIT \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/docs.github.com\/en\/site-policy\/github-terms\/github-terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">llama.cpp, The ggml authors<\/a><br><small>License: MIT \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 10<\/small><\/li><li><a href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\" rel=\"noopener noreferrer\">Max Vyaznikov, GPU Ark (Zenodo, doi:10.5281\/zenodo.20390790)<\/a><br><small>License: CC BY 4.0 \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/github.com\/meta-llama\/llama-models\/blob\/main\/models\/llama3_3\/LICENSE\" target=\"_blank\" rel=\"noopener noreferrer\">Meta Llama<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Oct 1, 2026 \u00b7 Source documents: 4<\/small><\/li><li><a href=\"https:\/\/pypi.org\/project\/mlx\/\" target=\"_blank\" rel=\"noopener noreferrer\">MLX, MLX Contributors<\/a><br><small>License: MIT \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/docs.github.com\/en\/site-policy\/github-terms\/github-terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">mlx-lm, Apple Inc.<\/a><br><small>License: MIT \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/pypi.org\/project\/mlx-lm\/\" target=\"_blank\" rel=\"noopener noreferrer\">mlx-lm, MLX Contributors<\/a><br><small>License: MIT \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/www.nvidia.com\/en-us\/geforce\/graphics-cards\/40-series\/rtx-4090\/\" target=\"_blank\" rel=\"noopener noreferrer\">NVIDIA<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 20, 2026 \u00b7 Source documents: 27<\/small><\/li><li><a href=\"https:\/\/arxiv.org\/html\/2505.09388\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen Team, Qwen3 Technical Report (arXiv:2505.09388)<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/qwenlm.github.io\/blog\/qwen3\/\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen Team<\/a><br><small>No data license stated \u00b7 Individual facts with source link \u00b7 Retrieved Sep 17, 2026 \u00b7 Source documents: 1<\/small><\/li><li><a href=\"https:\/\/docs.github.com\/en\/site-policy\/github-terms\/github-terms-of-service\" target=\"_blank\" rel=\"noopener noreferrer\">vLLM project<\/a><br><small>License: Apache 2.0 \u00b7 Used under license and terms of use \u00b7 Retrieved Sep 18, 2026 \u00b7 Source documents: 10<\/small><\/li><\/ul><\/details> <details class=\"tk-llm-method\" id=\"tk-llm-calculation\"><summary>Calculation method<\/summary><p>The memory calculation follows method memory-envelope-v1: <code>weights + 2*layers*kv_heads*head_dim*kv_bytes*context*concurrency + workspace<\/code><\/p><p>The KV cache gets a 5% allowance for memory management, the file size of the weights 3%. The runtime workspace is set at 1.0 to 3.0 GiB. A 10% reserve of device memory stays free.<\/p><p>The assumptions are based on <a href=\"https:\/\/docs.vllm.ai\/en\/latest\/configuration\/optimization\/\" target=\"_blank\" rel=\"noopener noreferrer\">the vLLM documentation on memory optimization<\/a> and are a planning assumption by torck, recorded on <time datetime=\"2026-09-17\">Sep 17, 2026<\/time>.<\/p><p class=\"tk-llm-method-example\">Worked example for Devstral-Small-2-24B-Instruct-2512 in Q4_K_M on GeForce RTX 4090 with 8,192 context tokens and concurrent requests 1. Weights 13.3 to 13.8 GiB, KV cache 1.3 GiB, runtime workspace 1.0 to 3.0 GiB. Together that is 15.6 to 18.1 GiB. 21.6 GiB are usable.<\/p><p>Computable means that the model variant has at least one configuration with a documented file size. This applies to 400 of 1,081 model variants in the catalogue. The catalogue lists the others with technical data but without a computable configuration yet.<\/p><p>Memory requirements are calculated, not measured. They consist of the model weights based on file size or quantization bit width, the KV cache for context and concurrent requests, and the runtime workspace. The selected reserve remains free.<\/p><p>Response speed is estimated, not measured. The upper bound is the memory bandwidth of the device divided by the bytes read for each generated token, i.e. the active parameters times bytes per weight plus the KV cache. The finder shows a range of 60 to 80 percent of this upper bound. This corridor is a rule of thumb, not a guarantee of accuracy. A number appears only for CUDA and Metal, for a single request and for models that fit entirely into device memory.<\/p><p>For runtime and device, each configuration names an evidence level, such as \u201cOfficially built\u201d or \u201cTarget architecture in the official build\u201d. An evidence level says nothing about speed or error-free output. If evidence is missing, the reason is listed under open points.<\/p><p>The estimate can be wrong in several cases. For MoE models, the calculation can be considerably too high. Long context, backend, driver and build, several concurrent requests, vLLM under load and offload change the real rate. For ROCm and SYCL, the finder shows no number because no verified reference measurement is available. Prompt processing before the first answer is not included.<\/p><\/details> <div class=\"tk-llm-mbar\" id=\"tk-llm-mbar\" hidden><span>Questions about your selection<\/span><a class=\"tk-llm-mbar-link\" href=\"\/en\/#contact-us\" aria-haspopup=\"dialog\" data-tk-action=\"contact-bar\">Contact<\/a><\/div> <script type=\"application\/json\" class=\"tk-llm-messages\">{\"error\":\"Invalid input. Check selections and number ranges.\",\"updating\":\"Calculating\\u2026\",\"compare_limit\":\"Select no more than four configurations.\",\"inputs_changed\":\"Selection changed. Recalculate the results.\",\"selection_count\":\"%d of 4 configurations selected\",\"load_more_status\":\"%1$s of %2$s shown\",\"load_more_announce\":\"%1$s of %2$s results shown.\"}<\/script>\n<\/section>\n<nav aria-label=\"More use cases and articles\"><ul class=\"tk-llm-entry-links\"><li><a href=\"https:\/\/www.torck.io\/en\/llm-finder\/dokumente-lokal\/\">Process documents with local LLMs<\/a><\/li><li><a href=\"https:\/\/www.torck.io\/en\/llm-finder\/coding-lokal\/\">Select a local coding LLM<\/a><\/li><li><a href=\"https:\/\/www.torck.io\/en\/llm-finder\/rtx-4090\/\">Plan LLM deployment on RTX 4090<\/a><\/li><li><a href=\"https:\/\/www.torck.io\/en\/llm-finder\/rag-deutsch\/\">Plan German RAG with local LLMs<\/a><\/li><\/ul><\/nav>\n","protected":false},"excerpt":{"rendered":"","protected":false},"author":4,"featured_media":0,"parent":5892,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_mbp_gutenberg_autopost":false,"footnotes":""},"class_list":["post-6012","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/pages\/6012","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/comments?post=6012"}],"version-history":[{"count":1,"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/pages\/6012\/revisions"}],"predecessor-version":[{"id":6016,"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/pages\/6012\/revisions\/6016"}],"up":[{"embeddable":true,"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/pages\/5892"}],"wp:attachment":[{"href":"https:\/\/www.torck.io\/en\/wp-json\/wp\/v2\/media?parent=6012"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}