# MODELLM > MODELLM ranks open-source language models against your requirements and sizes GPUs for your requests per second, with industry playbooks for banking, healthcare, education and more. ## Open-source models - [DeepSeek-V4-Flash-0731](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v4-flash-0731): DeepSeek-V4-Flash-0731 by deepseek-ai: a 304.2B-parameter mixture-of-experts open-weight model under the mit license. - [Qwen3.6-35B-A3B-NVFP4](https://modellm.outhentica.com/models/nvidia-qwen3-6-35b-a3b-nvfp4): Qwen3.6-35B-A3B-NVFP4 by nvidia: a 18.7B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling and vision. - [Qwen3-8B](https://modellm.outhentica.com/models/qwen-qwen3-8b): Qwen3-8B by Qwen: a 8.2B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [OTel-2.0-LLM-31B-IT](https://modellm.outhentica.com/models/farbodtavakkoli-otel-2-0-llm-31b-it): OTel-2.0-LLM-31B-IT by farbodtavakkoli: a 31.3B-parameter dense open-weight model under the apache-2.0 license and vision. - [Llama 3.3 70B Instruct](https://modellm.outhentica.com/models/llama-3-3-70b-instruct): Meta's 70B open-weight model with strong reasoning and 128K context; needs roughly 45 GB at INT4 plus KV cache. - [gpt-oss-20b](https://modellm.outhentica.com/models/openai-gpt-oss-20b): gpt-oss-20b by openai: a 20.9B-parameter mixture-of-experts open-weight model under the apache-2.0 license. - [dolphin-2.9.1-yi-1.5-34b](https://modellm.outhentica.com/models/dphn-dolphin-2-9-1-yi-1-5-34b): dolphin-2.9.1-yi-1.5-34b by dphn: a 34.4B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen2.5-3B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-3b-instruct): Qwen2.5-3B-Instruct by Qwen: a 3.1B-parameter dense open-weight model under the other license with tool calling. - [Kimi-K3-DSpark](https://modellm.outhentica.com/models/radixark-kimi-k3-dspark): Kimi-K3-DSpark by RadixArk: a 2.2B-parameter dense open-weight model under the unknown license. - [NVIDIA-Nemotron-3-Nano-4B-BF16](https://modellm.outhentica.com/models/nvidia-nvidia-nemotron-3-nano-4b-bf16): NVIDIA-Nemotron-3-Nano-4B-BF16 by nvidia: a 4B-parameter dense open-weight model under the other license with tool calling. - [Qwen2.5-VL-32B-Instruct](https://modellm.outhentica.com/models/qwen2-5-vl-32b): An Apache-2.0 vision-language model that reads scanned documents, forms and IDs; the usual pick when OCR-free extraction is required. - [DeepSeek-V3.2](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v3-2): DeepSeek-V3.2 by deepseek-ai: a 685.4B-parameter mixture-of-experts open-weight model under the mit license. - [Gemma-4-26B-A4B-NVFP4](https://modellm.outhentica.com/models/nvidia-gemma-4-26b-a4b-nvfp4): Gemma-4-26B-A4B-NVFP4 by nvidia: a 14.4B-parameter mixture-of-experts open-weight model under the apache-2.0 license and vision. - [Qwen3-30B-A3B](https://modellm.outhentica.com/models/qwen-qwen3-30b-a3b): Qwen3-30B-A3B by Qwen: a 30.5B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling. - [Qwen3.5-122B-A10B-NVFP4](https://modellm.outhentica.com/models/nvidia-qwen3-5-122b-a10b-nvfp4): Qwen3.5-122B-A10B-NVFP4 by nvidia: a 64.6B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling and vision. - [Qwen2.5-Coder-14B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-coder-14b-instruct): Qwen2.5-Coder-14B-Instruct by Qwen: a 14.8B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen3-4B-Base](https://modellm.outhentica.com/models/qwen-qwen3-4b-base): Qwen3-4B-Base by Qwen: a 4B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen3-1.7B-Base](https://modellm.outhentica.com/models/qwen-qwen3-1-7b-base): Qwen3-1.7B-Base by Qwen: a 1.7B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Gemma-4-31B-IT-NVFP4](https://modellm.outhentica.com/models/nvidia-gemma-4-31b-it-nvfp4): Gemma-4-31B-IT-NVFP4 by nvidia: a 20.9B-parameter dense open-weight model under the other license and vision. - [DeepSeek-V4-Flash](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v4-flash): DeepSeek-V4-Flash by deepseek-ai: a 290.9B-parameter mixture-of-experts open-weight model under the mit license. - [Qwen2.5-1.5B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-1-5b-instruct): Qwen2.5-1.5B-Instruct by Qwen: a 1.5B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [gpt-oss-120b](https://modellm.outhentica.com/models/openai-gpt-oss-120b): gpt-oss-120b by openai: a 116.8B-parameter mixture-of-experts open-weight model under the apache-2.0 license. - [Qwen3-32B](https://modellm.outhentica.com/models/qwen-qwen3-32b): Qwen3-32B by Qwen: a 32.8B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen2.5-Coder-7B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-coder-7b-instruct): Qwen2.5-Coder-7B-Instruct by Qwen: a 7.6B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen2.5-14B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-14b-instruct): Qwen2.5-14B-Instruct by Qwen: a 14.8B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Mistral Small 3 (24B)](https://modellm.outhentica.com/models/mistral-small-3-24b): A fast 24B Apache-2.0 model tuned for low latency; a single 48 GB GPU serves moderate traffic at INT4. - [Mistral-7B-Instruct-v0.2](https://modellm.outhentica.com/models/mistralai-mistral-7b-instruct-v0-2): Mistral-7B-Instruct-v0.2 by mistralai: a 7.2B-parameter dense open-weight model under the apache-2.0 license. - [TinyLlama-1.1B-Chat-v1.0](https://modellm.outhentica.com/models/tinyllama-tinyllama-1-1b-chat-v1-0): TinyLlama-1.1B-Chat-v1.0 by TinyLlama: a 1.1B-parameter dense open-weight model under the apache-2.0 license. - [PowerMoE-3b](https://modellm.outhentica.com/models/ibm-research-powermoe-3b): PowerMoE-3b by ibm-research: a 3.4B-parameter mixture-of-experts open-weight model under the apache-2.0 license. - [NVIDIA-Nemotron-3-Super-120B-A12B-BF16](https://modellm.outhentica.com/models/nvidia-nvidia-nemotron-3-super-120b-a12b-bf16): NVIDIA-Nemotron-3-Super-120B-A12B-BF16 by nvidia: a 123.6B-parameter mixture-of-experts open-weight model under the other license. - [Qwen3.8-27B-OBLITERATED](https://modellm.outhentica.com/models/obliteratus-qwen3-8-27b-obliterated): Qwen3.8-27B-OBLITERATED by OBLITERATUS: a 27.8B-parameter dense open-weight model under the apache-2.0 license and vision. - [Qwen-72B](https://modellm.outhentica.com/models/qwen-qwen-72b): Qwen-72B by Qwen: a 72.3B-parameter dense open-weight model under the other license. - [Qwen2.5-7B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-7b-instruct): Qwen2.5-7B-Instruct by Qwen: a 7.6B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen3-4B](https://modellm.outhentica.com/models/qwen-qwen3-4b): Qwen3-4B by Qwen: a 4B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Ornith-1.5-35B-A3B-NVFP4](https://modellm.outhentica.com/models/ornith-ai-ornith-1-5-35b-a3b-nvfp4): Ornith-1.5-35B-A3B-NVFP4 by ornith-ai: a 19.5B-parameter mixture-of-experts open-weight model under the mit license with tool calling and vision. - [Qwen3-Coder-30B-A3B-Instruct-FP8](https://modellm.outhentica.com/models/qwen-qwen3-coder-30b-a3b-instruct-fp8): Qwen3-Coder-30B-A3B-Instruct-FP8 by Qwen: a 30.5B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling. - [Qwen3-1.7B](https://modellm.outhentica.com/models/qwen-qwen3-1-7b): Qwen3-1.7B by Qwen: a 2B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen3-14B](https://modellm.outhentica.com/models/qwen-qwen3-14b): Qwen3-14B by Qwen: a 14.8B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [Qwen2.5-32B-Instruct](https://modellm.outhentica.com/models/qwen2-5-32b-instruct): A 32B dense open-weight model under Apache-2.0 with 128K context and broad multilingual coverage; fits on two 48 GB GPUs at INT4. - [Qwen3-4B-Instruct-2507](https://modellm.outhentica.com/models/qwen-qwen3-4b-instruct-2507): Qwen3-4B-Instruct-2507 by Qwen: a 4B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [GLM-4.7-Flash](https://modellm.outhentica.com/models/zai-org-glm-4-7-flash): GLM-4.7-Flash by zai-org: a 31.2B-parameter mixture-of-experts open-weight model under the mit license. - [MiniMax-M2.7](https://modellm.outhentica.com/models/minimaxai-minimax-m2-7): MiniMax-M2.7 by MiniMaxAI: a 228.7B-parameter mixture-of-experts open-weight model under the other license. - [Qwen3-Coder-Next-FP8](https://modellm.outhentica.com/models/qwen-qwen3-coder-next-fp8): Qwen3-Coder-Next-FP8 by Qwen: a 79.7B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling. - [Qwen2.5-Coder-32B-Instruct](https://modellm.outhentica.com/models/qwen-qwen2-5-coder-32b-instruct): Qwen2.5-Coder-32B-Instruct by Qwen: a 32.8B-parameter dense open-weight model under the apache-2.0 license with tool calling. - [DeepSeek-V3](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v3): DeepSeek-V3 by deepseek-ai: a 684.5B-parameter mixture-of-experts open-weight model under the unknown license with tool calling. - [NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4](https://modellm.outhentica.com/models/nvidia-nvidia-nemotron-3-5-lightning-30b-a3b-nvfp4): NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 by nvidia: a 17.8B-parameter mixture-of-experts open-weight model under the other license. - [Bonsai-27B-mlx-1bit](https://modellm.outhentica.com/models/prism-ml-bonsai-27b-mlx-1bit): Bonsai-27B-mlx-1bit by prism-ml: a 1.7B-parameter dense open-weight model under the apache-2.0 license and vision. - [Ternary-Bonsai-27B-mlx-2bit](https://modellm.outhentica.com/models/prism-ml-ternary-bonsai-27b-mlx-2bit): Ternary-Bonsai-27B-mlx-2bit by prism-ml: a 27.4B-parameter dense open-weight model under the apache-2.0 license and vision. - [Llama-3.1-8B-Instruct-4bit](https://modellm.outhentica.com/models/mlx-community-llama-3-1-8b-instruct-4bit): Llama-3.1-8B-Instruct-4bit by mlx-community: a 8B-parameter dense open-weight model under the llama3.1 license with tool calling. - [DeepSeek-V4-Flash-DSpark](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v4-flash-dspark): DeepSeek-V4-Flash-DSpark by deepseek-ai: a 165.3B-parameter mixture-of-experts open-weight model under the mit license. - [DeepSeek-V3-0324](https://modellm.outhentica.com/models/deepseek-ai-deepseek-v3-0324): DeepSeek-V3-0324 by deepseek-ai: a 684.5B-parameter mixture-of-experts open-weight model under the mit license with tool calling. - [JiRackUltra_14b](https://modellm.outhentica.com/models/cmsmanhattan-jirackultra-14b): JiRackUltra_14b by CMSManhattan: a 14.8B-parameter dense open-weight model under the mit license. - [GLM-5.2-FP8](https://modellm.outhentica.com/models/zai-org-glm-5-2-fp8): GLM-5.2-FP8 by zai-org: a 753.3B-parameter mixture-of-experts open-weight model under the mit license. - [GLM-5.2](https://modellm.outhentica.com/models/zai-org-glm-5-2): GLM-5.2 by zai-org: a 753.3B-parameter mixture-of-experts open-weight model under the mit license. - [GLM-5.3](https://modellm.outhentica.com/models/zai-org-glm-5-3): GLM-5.3 by zai-org: a 753.3B-parameter mixture-of-experts open-weight model under the other license. - [Qwen3-4B-Instruct-2507-FP8](https://modellm.outhentica.com/models/qwen-qwen3-4b-instruct-2507-fp8): Qwen3-4B-Instruct-2507-FP8 by Qwen: a 4.4B-parameter dense open-weight model under the apache-2.0 license with tool calling. ## Industry playbooks - [Banking](https://modellm.outhentica.com/industries/banking): Banks run customer service, KYC extraction, credit summarisation and policy search on open models kept inside the data centre. - [Fintech & payments](https://modellm.outhentica.com/industries/fintech): Wallets, lenders and payment processors use open models for support automation, dispute handling, merchant onboarding and fraud-analyst tooling. - [Insurance](https://modellm.outhentica.com/industries/insurance): Insurers use open models for claims intake, policy servicing, underwriting summaries and fraud triage with a full audit trail. - [Healthcare & hospitals](https://modellm.outhentica.com/industries/healthcare): Hospitals use open models for discharge summaries, coding, patient-message triage and clinical search with PHI kept on-premises and clinicians in the loop. - [Pharma & life sciences](https://modellm.outhentica.com/industries/pharma): Pharma teams use open models for literature review, regulatory writing, pharmacovigilance case processing and lab-protocol assistance under GxP controls. - [Education](https://modellm.outhentica.com/industries/education): Schools and universities use open models for tutoring, grading support, course-material generation and admin help, with student-data protection built in. - [Government & public sector](https://modellm.outhentica.com/industries/government): Agencies use open models for citizen services, case-work summarisation, policy research and records management under data-sovereignty rules. - [Telecom](https://modellm.outhentica.com/industries/telecom): Operators use open models for very high-volume customer support, network-operations copilots, billing dispute handling and CDR-safe analytics. - [Retail & e-commerce](https://modellm.outhentica.com/industries/retail): Retailers use open models for product search, catalogue enrichment, customer support, returns handling and merchandising copy at high volume. - [Manufacturing](https://modellm.outhentica.com/industries/manufacturing): Manufacturers use open models for maintenance copilots, quality-incident analysis, SOP search, supplier documents and shop-floor multilingual support. - [Energy & utilities](https://modellm.outhentica.com/industries/energy): Utilities and energy companies use open models for outage communication, asset-maintenance copilots, regulatory filings and field-worker support. - [Legal services](https://modellm.outhentica.com/industries/legal): Law firms and legal departments use open models for contract review, due diligence, research memos and e-discovery with privilege protected on-premises. - [Real estate & property](https://modellm.outhentica.com/industries/real-estate): Developers, brokers and property managers use open models for listing content, lease abstraction, tenant support and valuation research. - [Logistics & supply chain](https://modellm.outhentica.com/industries/logistics): Carriers, forwarders and 3PLs use open models for shipment-status support, customs documentation, exception handling and dispatcher copilots. - [Media & entertainment](https://modellm.outhentica.com/industries/media): Publishers, broadcasters and studios use open models for content tagging, summarisation, localisation, archive search and audience support. - [Travel & hospitality](https://modellm.outhentica.com/industries/hospitality): Hotels, airlines and travel platforms use open models for multilingual guest service, booking changes, itinerary planning and review response. - [HR & recruiting](https://modellm.outhentica.com/industries/hr): HR teams and staffing firms use open models for candidate screening support, policy Q&A, job-description drafting and employee helpdesks, with bias controls. - [IT & software](https://modellm.outhentica.com/industries/software): Software companies and IT departments use open models for coding assistants, internal knowledge search, incident response, and customer technical support. - [Agriculture & food](https://modellm.outhentica.com/industries/agriculture): Agribusinesses and food producers use open models for agronomy advice, traceability documents, quality reports and farmer-facing multilingual support. - [Automotive & mobility](https://modellm.outhentica.com/industries/automotive): OEMs, dealers and mobility providers use open models for in-vehicle assistants, dealer service support, warranty-claim analysis and engineering documentation. - [Other / general](https://modellm.outhentica.com/industries/general): For sectors without a dedicated playbook: the generic questionnaire with no industry-specific compliance checks. ## Blog - [How to size GPUs for an open-source LLM: the three numbers that matter](https://modellm.outhentica.com/blog/how-to-size-gpus-for-an-open-llm): Weights are the easy part. KV cache multiplied by concurrency is what actually decides how many GPUs you need — and it's the number most first estimates leave out. - [Open LLM licenses explained: Apache-2.0, MIT, Llama, Gemma and the research-only crowd](https://modellm.outhentica.com/blog/open-llm-licenses-explained): Not every open-weight model is open for commercial use in the same way. Here is what the common licenses allow, what they restrict, and what procurement should check. - [Dense vs mixture-of-experts: what it means for your hardware bill](https://modellm.outhentica.com/blog/dense-vs-moe-models): MoE models activate a fraction of their parameters per token. That changes compute and latency — but not memory, which is where naive sizing goes wrong. - [vLLM, SGLang, TensorRT-LLM or llama.cpp: choosing an inference server](https://modellm.outhentica.com/blog/choosing-an-inference-server): The model is half the decision. The server decides your throughput, your latency profile and how much of your GPU you actually use. - [On-premises vs private cloud for regulated LLM deployments](https://modellm.outhentica.com/blog/on-prem-vs-private-cloud-for-regulated-llm): Banks, hospitals and governments all ask the same question. The answer depends on three things: data classification, who holds the keys, and how fast you need to scale. - [Quantization in practice: AWQ, GPTQ, GGUF and FP8](https://modellm.outhentica.com/blog/quantization-awq-gptq-gguf-fp8): Quantized weights are how most open models reach production. Here is what each format is for, what it costs in quality, and how to verify it on your own data. ## Adoption stories - [Open models in production](https://modellm.outhentica.com/stories): Publicly reported enterprise deployments of open-weight models, with sources ## Other - [FAQ](https://modellm.outhentica.com/faq): Common questions about choosing and hosting open LLMs