One assessment.
Every open model.
Rank 56 open-weight LLMs against your requirements with hard filters first and a weighted score second — every pick explains why.
{
"profile": { "tasks": ["rag","chat"], "deployment": "on_prem" },
"best_fit": { "model": "Qwen2.5-32B-Instruct", "score": 91 },
"sizing": { "total_gb": 73, "gpus": "2 × L40S 48 GB" }
}Infrastructure, not just a list.
Replace spreadsheet comparisons with a single pipeline from requirements to a GPU bill of materials.
Unified catalog
Every model row carries architecture, license, features and Hub metrics, synced automatically.
Industry playbooks
Sector-specific use cases with risk tiers set filters and weights for you.
Explainable ranking
Hard filters first, then a weighted score. Every pick shows why, and what it trades off.
✓ 128K context covers 32K
✓ strong on RAG, extraction
Request-to-GPU sizing
Peak requests per second and token lengths become memory, then real GPU configurations.
The problem with choosing a model by hand
You want to ship an assistant. Instead you're maintaining spreadsheets.
Fragmented sources
Licenses on the Hub, benchmarks in papers, memory numbers in blog posts. Nothing lines up.
Compliance overhead
Every regulated sector has data rules that decide where a model may run — usually discovered late.
Memory surprises
Teams size on weights and forget the KV cache. The first load test is where the GPU order gets doubled.
Constant churn
New models weekly, inference stacks monthly. Last quarter's shortlist is already stale.
How it works
Requirements in, deployment plan out.
Requirement profile
Both intake paths produce one structured object you confirm before anything is recommended.
tasks: ["rag", "chat"]
languages: ["en", "ur"]
contextTokens: 32000
deployment: "on_prem"
traffic: { peakRps: 10, in: 2000, out: 400 }Availability check
Hard filters remove models that can't legally or technically do the job before any scoring.
Ranked options
Best fit, highest quality and cost-optimized — never a single black-box answer.
Request lifecycle
Latency targets feed the sizing: prefill, decode speed and batch size decide concurrency.
From profile to hardware
One profile fans out into a model pick and a GPU configuration.
The open LLM landscape, live.
Numbers below come from the catalog and refresh with every Hugging Face sync.
- 56
- open models tracked
- 156,617,546
- combined Hub downloads
- 23
- mixture-of-experts
- 33
- dense
- 28
- with tool calling
- 10
- with vision
Qwen3-8B by Qwen: a 8.2B-parameter dense open-weight model under the apache-2.0 license with tool calling.
Qwen2.5-7B-Instruct by Qwen: a 7.6B-parameter dense open-weight model under the apache-2.0 license with tool calling.
Qwen3.6-35B-A3B-NVFP4 by nvidia: a 18.7B-parameter mixture-of-experts open-weight model under the apache-2.0 license with tool calling and vision.
Qwen2.5-1.5B-Instruct by Qwen: a 1.5B-parameter dense open-weight model under the apache-2.0 license with tool calling.
Qwen3-4B by Qwen: a 4B-parameter dense open-weight model under the apache-2.0 license with tool calling.
OTel-2.0-LLM-31B-IT by farbodtavakkoli: a 31.3B-parameter dense open-weight model under the apache-2.0 license and vision.
Compliance, handled.
Every playbook encodes the data rules, regulations and systems of its sector, so the recommendation is deployable, not just clever.
All 21 industries- Agriculture & food
- Automotive & mobility
- Banking
- Education
- Energy & utilities
- Fintech & payments
- Government & public sector
- HR & recruiting
- Healthcare & hospitals
- IT & software
- Insurance
- Legal services
- Logistics & supply chain
- Manufacturing
- Media & entertainment
- Other / general
- Pharma & life sciences
- Real estate & property
- Retail & e-commerce
- Telecom
- Travel & hospitality
Live updates
All updates →Hugging Face · HF Papers · arXiv · GitHub — refreshed every 15 minutes
- 2026-09-20Hugging Face🔥 Trending on Hugging Face: Bonsai-2-27B-Ternary-CRACK-GGUF
- 2026-09-20Hugging Face🔥 Trending on Hugging Face: Qwen3.8-35B-A3B-Distill-GGUF
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen3-4B-Instruct-2507-FP8
- 2026-09-19Hugging Face🆕 New on the Hub: GLM-5.3
- 2026-09-19Hugging Face🆕 New on the Hub: GLM-5.2
- 2026-09-19Hugging Face🆕 New on the Hub: GLM-5.2-FP8
- 2026-09-19Hugging Face🆕 New on the Hub: JiRackUltra_14b
- 2026-09-19Hugging Face🆕 New on the Hub: DeepSeek-V3-0324
- 2026-09-19Hugging Face🆕 New on the Hub: DeepSeek-V4-Flash-DSpark
- 2026-09-19Hugging Face🆕 New on the Hub: Llama-3.1-8B-Instruct-4bit
- 2026-09-19Hugging Face🆕 New on the Hub: Ternary-Bonsai-27B-mlx-2bit
- 2026-09-19Hugging Face🆕 New on the Hub: Bonsai-27B-mlx-1bit
- 2026-09-19Hugging Face🆕 New on the Hub: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen3-Coder-30B-A3B-Instruct-FP8
- 2026-09-19Hugging Face🆕 New on the Hub: DeepSeek-V3
- 2026-09-19Hugging Face🆕 New on the Hub: Ornith-1.5-35B-A3B-NVFP4
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen3.8-27B-OBLITERATED
- 2026-09-19Hugging Face🆕 New on the Hub: NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- 2026-09-19Hugging Face🆕 New on the Hub: PowerMoE-3b
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen2.5-Coder-32B-Instruct
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen3-Coder-Next-FP8
- 2026-09-19Hugging Face🆕 New on the Hub: Qwen3.5-122B-A10B-NVFP4
- 2026-09-19Hugging Face🆕 New on the Hub: TinyLlama-1.1B-Chat-v1.0
- 2026-09-19Hugging Face🆕 New on the Hub: MiniMax-M2.7
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: Ternary-Bonsai-2-27B-gguf
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: MiniCPM5-2B
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: Qwen-2.5-1B-RLCD
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: needle3
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: Nex-N2.5-mini
- 2026-09-19Hugging Face🔥 Trending on Hugging Face: K2-Horizon-7B
Guides
All posts →- 2026-09-19
📐 How to size GPUs for an open-source LLM: the three numbers that matter
Weights are the easy part. KV cache multiplied by concurrency is what actually decides how many GPUs you need — and it's the number most first estimates leave out.
- 2026-09-19
⚖️ Open LLM licenses explained: Apache-2.0, MIT, Llama, Gemma and the research-only crowd
Not every open-weight model is open for commercial use in the same way. Here is what the common licenses allow, what they restrict, and what procurement should check.
- 2026-09-19
🧩 Dense vs mixture-of-experts: what it means for your hardware bill
MoE models activate a fraction of their parameters per token. That changes compute and latency — but not memory, which is where naive sizing goes wrong.
- 2026-09-19
⚙️ vLLM, SGLang, TensorRT-LLM or llama.cpp: choosing an inference server
The model is half the decision. The server decides your throughput, your latency profile and how much of your GPU you actually use.
- 2026-09-19
🏛️ On-premises vs private cloud for regulated LLM deployments
Banks, hospitals and governments all ask the same question. The answer depends on three things: data classification, who holds the keys, and how fast you need to scale.
- 2026-09-19
🗜️ Quantization in practice: AWQ, GPTQ, GGUF and FP8
Quantized weights are how most open models reach production. Here is what each format is for, what it costs in quality, and how to verify it on your own data.
Open models in production
Publicly reported deployments across banking, healthcare, government and logistics. Each card links to its source.
- GermanyHealthcare & banking
Confidential-computing LLM service for regulated customers
Edgeless Systems · Privatemode AIBuilt an end-to-end encrypted LLM service for healthcare and banking clients whose data rules ruled out ordinary cloud AI. Chose Llama because it was open, multilingual and self-hostable, and ran a quantised INT4 build to balance quality against compute.
Llama 3.1 70B → Llama 3.3 70B (AWQ INT4)Source: Meta — Llama case study ↗ - Japan / globalFinancial services
Firm-wide generative AI on open models
NomuraThe investment bank set out to make generative AI available across the firm and standardised on Llama models, citing faster innovation, transparency, bias guardrails and strong results on summarisation and code generation.
Llama (Meta)Source: AWS customer case study ↗ - United StatesFintech & payments
Customer support systems built on an open model
Block · Cash AppBlock integrated Llama into the support systems behind Cash App. Because the weights are open, the team can experiment and customise per use case while keeping customer data private.
Llama (Meta)Source: Meta AI blog ↗ - BrazilHealthcare & hospitals
Automated hospital discharge summaries
Instituto de Inteligência Artificial na Saúde · NoHarmThe NoHarm Summary Discharge project drafts and improves patient discharge summaries to cut administrative load, reduce errors and give patients complete post-hospital instructions — one of the highest-value, highest-risk hospital use cases in the playbook.
Llama (Meta)Source: Meta — Llama AI innovation ↗ - United Arab EmiratesHealthcare & research
Bilingual Arabic–English biomedical assistant
MBZUAI · BiMediX2A bilingual expert system providing virtual medical advice and diagnostic support in English and Arabic, including medical-image queries and spoken interaction — an example of open models serving languages that closed APIs cover poorly.
Llama (Meta)Seamless-M4TSource: Meta — Llama AI innovation ↗ - US · EU · Japan · Korea · NATO-associatedGovernment & public sector
Open models for sovereign and national-security use
US, EU and allied governmentsMeta extended Llama access to US agencies and contractors and later to governments including France, Germany, Italy, Japan and South Korea, noting that open weights can be downloaded and deployed in secure environments at various classification levels without sending data to third-party providers.
Llama (Meta)Source: Engadget (via Yahoo Tech) ↗
Pricing
Run assessments, browse the catalog, read the feed.
- ✓ Unlimited assessments
- ✓ Full model catalog
- ✓ Industry playbooks
- ✓ Live updates feed
Team workspaces, private catalogs and procurement-ready outputs.
- ✓ Saved projects & sharing
- ✓ Private model catalog
- ✓ Exportable vLLM / Helm configs
- ✓ Compliance summary reports
- ✓ Priority support
FAQ
What is MODELLM?+
MODELLM is a selection and sizing platform for open-source language models. It ranks open-weight models against your requirements, then estimates the GPU hardware needed for your requests per second.
Where does the model data come from?+
Specifications are synced from each model's Hugging Face repository (config, license, languages, downloads) and GitHub. Editorial task scores and summaries are curated separately.
How accurate is the hardware sizing?+
It is an analytical first-order estimate: weights plus KV cache times concurrency plus headroom. Expect real throughput to land within roughly ±30 %; validate with a load test before purchasing.
Does it support on-premises and air-gapped deployments?+
Yes. Industry playbooks ask where data may be processed and filter recommendations to models and configurations that can run inside your own data centre.
Which industries have playbooks?+
Banking, fintech, insurance, healthcare, pharma, education, government, telecom, retail, manufacturing, energy, legal, real estate, logistics, media, hospitality, HR, software, agriculture and automotive, plus a general path.
Is MODELLM free?+
Running an assessment is free. Enterprise plans add saved projects, team workspaces, private catalogs and exportable deployment configurations.