Updates
New and trending models on Hugging Face, the latest LLM research from arXiv and HF Daily Papers, and releases of the major open inference stacks on GitHub.
Last refreshed 2026-09-19 23:04 UTC · refreshes automatically every 15 minutes
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-4B-Instruct-2507-FP8
Qwen published Qwen3-4B-Instruct-2507-FP8, a 4.4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
zai-org published GLM-5.3, a 753.3B moe model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
zai-org published GLM-5.2, a 753.3B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
zai-org published GLM-5.2-FP8, a 753.3B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: JiRackUltra_14b
CMSManhattan published JiRackUltra_14b, a 14.8B dense model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: DeepSeek-V3-0324
deepseek-ai published DeepSeek-V3-0324, a 684.5B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: DeepSeek-V4-Flash-DSpark
deepseek-ai published DeepSeek-V4-Flash-DSpark, a 165.3B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Llama-3.1-8B-Instruct-4bit
mlx-community published Llama-3.1-8B-Instruct-4bit, a 8B dense model under llama3.1. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Ternary-Bonsai-27B-mlx-2bit
prism-ml published Ternary-Bonsai-27B-mlx-2bit, a 27.4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Bonsai-27B-mlx-1bit
prism-ml published Bonsai-27B-mlx-1bit, a 1.7B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia published NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4, a 17.8B moe model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-Coder-30B-A3B-Instruct-FP8
Qwen published Qwen3-Coder-30B-A3B-Instruct-FP8, a 30.5B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
deepseek-ai published DeepSeek-V3, a 684.5B moe model under unknown. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Ornith-1.5-35B-A3B-NVFP4
ornith-ai published Ornith-1.5-35B-A3B-NVFP4, a 19.5B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3.8-27B-OBLITERATED
OBLITERATUS published Qwen3.8-27B-OBLITERATED, a 27.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia published NVIDIA-Nemotron-3-Super-120B-A12B-BF16, a 123.6B moe model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
ibm-research published PowerMoE-3b, a 3.4B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-Coder-32B-Instruct
Qwen published Qwen2.5-Coder-32B-Instruct, a 32.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-Coder-Next-FP8
Qwen published Qwen3-Coder-Next-FP8, a 79.7B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3.5-122B-A10B-NVFP4
nvidia published Qwen3.5-122B-A10B-NVFP4, a 64.6B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: TinyLlama-1.1B-Chat-v1.0
TinyLlama published TinyLlama-1.1B-Chat-v1.0, a 1.1B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: MiniMax-M2.7
MiniMaxAI published MiniMax-M2.7, a 228.7B moe model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: NeoHorse-1-9B
TokenRhythm/NeoHorse-1-9B is trending with 11,692 downloads and 944 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Xing4.0-29B-A4B
XingChen-AGI/Xing4.0-29B-A4B is trending with 7,278 downloads and 619 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: penclaw-GLM-5.3-abliterated
audnai/penclaw-GLM-5.3-abliterated is trending with 626 downloads and 164 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Nex-N2.5-mini
nex-agi/Nex-N2.5-mini is trending with 8,239 downloads and 838 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: needle3
Cactus-Compute/needle3 is trending with 24,420 downloads and 94 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Ternary-Bonsai-2-27B-mlx-2bit
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit is trending with 23,111 downloads and 238 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Edge0-35B-A3B-preview
Edge0/Edge0-35B-A3B-preview is trending with 68,403 downloads and 3,470 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Llama-3.1-8B-Instruct
meta-llama/Llama-3.1-8B-Instruct is trending with 5,919,746 downloads and 7,740 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Spark-X2.5-4B
XHToken/Spark-X2.5-4B is trending with 30,215 downloads and 1,285 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: DeepSeek-R1
deepseek-ai/DeepSeek-R1 is trending with 760,320 downloads and 14,266 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: gpt2
openai-community/gpt2 is trending with 15,362,811 downloads and 4,150 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: K2-Horizon-7B
IFM/K2-Horizon-7B is trending with 14,431 downloads and 213 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: MiniCPM5-2B
openbmb/MiniCPM5-2B is trending with 389,555 downloads and 1,589 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Qwen-2.5-1B-RLCD
harshatheg/Qwen-2.5-1B-RLCD is trending with 0 downloads and 424 likes.
- 2026-09-19·Hugging Face
🔥 Trending on Hugging Face: Ternary-Bonsai-2-27B-gguf
prism-ml/Ternary-Bonsai-2-27B-gguf is trending with 1,516,960 downloads and 1,191 likes.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-1.7B-Base
Qwen published Qwen3-1.7B-Base, a 1.7B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-4B-Base
Qwen published Qwen3-4B-Base, a 4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: DeepSeek-V4-Flash
deepseek-ai published DeepSeek-V4-Flash, a 290.9B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Gemma-4-31B-IT-NVFP4
nvidia published Gemma-4-31B-IT-NVFP4, a 20.9B dense model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Mistral-7B-Instruct-v0.2
mistralai published Mistral-7B-Instruct-v0.2, a 7.2B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-30B-A3B
Qwen published Qwen3-30B-A3B, a 30.5B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Gemma-4-26B-A4B-NVFP4
nvidia published Gemma-4-26B-A4B-NVFP4, a 14.4B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-Coder-14B-Instruct
Qwen published Qwen2.5-Coder-14B-Instruct, a 14.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: GLM-4.7-Flash
zai-org published GLM-4.7-Flash, a 31.2B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen3-14B, a 14.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-14B-Instruct
Qwen published Qwen2.5-14B-Instruct, a 14.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: DeepSeek-V3.2
deepseek-ai published DeepSeek-V3.2, a 685.4B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-Coder-7B-Instruct
Qwen published Qwen2.5-Coder-7B-Instruct, a 7.6B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: NVIDIA-Nemotron-3-Nano-4B-BF16
nvidia published NVIDIA-Nemotron-3-Nano-4B-BF16, a 4B dense model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Kimi-K3-DSpark
RadixArk published Kimi-K3-DSpark, a 2.2B dense model under unknown. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen3-1.7B, a 2B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3-4B-Instruct-2507
Qwen published Qwen3-4B-Instruct-2507, a 4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen-72B, a 72.3B dense model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: DeepSeek-V4-Flash-0731
deepseek-ai published DeepSeek-V4-Flash-0731, a 304.2B moe model under mit. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: dolphin-2.9.1-yi-1.5-34b
dphn published dolphin-2.9.1-yi-1.5-34b, a 34.4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen3-32B, a 32.8B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-3B-Instruct
Qwen published Qwen2.5-3B-Instruct, a 3.1B dense model under other. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: gpt-oss-120b
openai published gpt-oss-120b, a 116.8B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
openai published gpt-oss-20b, a 20.9B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: OTel-2.0-LLM-31B-IT
farbodtavakkoli published OTel-2.0-LLM-31B-IT, a 31.3B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen3-4B, a 4B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-1.5B-Instruct
Qwen published Qwen2.5-1.5B-Instruct, a 1.5B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen3.6-35B-A3B-NVFP4
nvidia published Qwen3.6-35B-A3B-NVFP4, a 18.7B moe model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
🆕 New on the Hub: Qwen2.5-7B-Instruct
Qwen published Qwen2.5-7B-Instruct, a 7.6B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-19·Hugging Face
Qwen published Qwen3-8B, a 8.2B dense model under apache-2.0. Imported as a draft for review.
- 2026-09-18·GitHub
Highlights 713 PRs from 237 contributors. New models in this release (see the cookbook(https://docs.sglang.io/cookbook) for all supported models): | Model | Type | PRs | Cookbook | |---|---|---|---| | GLM-5.3-Flash | Aut…
- 2026-09-17·arXiv
🧪 Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorizat…
- 2026-09-17·arXiv
🧪 Score Centering Stabilizes Off-policy Reinforcement Learning
Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely e…
- 2026-09-17·arXiv
🧪 An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effe…
- 2026-09-17·arXiv
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit d…
- 2026-09-17·arXiv
🧪 Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adop…
- 2026-09-17·arXiv
This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable …
- 2026-09-17·arXiv
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, …
- 2026-09-17·arXiv
🧪 Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-st…
- 2026-09-17·arXiv
🧪 WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution
Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also…
- 2026-09-17·arXiv
🧪 SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment
Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional…
- 2026-09-17·arXiv
Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic …
- 2026-09-17·arXiv
🧪 Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language …
- 2026-09-17·arXiv
🧪 An Analysis of Training-Free Self-Reported Confidence in Language Models
Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbaliz…
- 2026-09-17·arXiv
🧪 Parallelism, critical windows, and separations among diffusion language models
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward p…
- 2026-09-17·arXiv
🧪 Edustories: A Collection of Real-world Case Studies from Classroom Practices
Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective cl…
- 2026-09-17·HF Papers
📄 FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation a…
- 2026-09-17·HF Papers
📄 RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with…
- 2026-09-17·HF Papers
📄 Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as predic…
- 2026-09-17·HF Papers
📄 What Does Privileged Information Add to On-Policy Self-Distillation?
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, bu…
- 2026-09-17·HF Papers
📄 WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clea…
- 2026-09-17·HF Papers
📄 VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored…
- 2026-09-17·HF Papers
📄 When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post…
- 2026-09-17·HF Papers
📄 When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length pena…
- 2026-09-17·HF Papers
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval qua…
- 2026-09-16·HF Papers
📄 Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context underst…
- 2026-09-16·HF Papers
📄 PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's…
- 2026-09-16·HF Papers
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candida…
- 2026-09-15·GitHub
What's Changed - Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. - Added ollama://apps to open the desktop app…
- 2026-09-15·HF Papers
📄 Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before execution. Recent agent-skill frameworks encapsulate…
- 2026-09-15·HF Papers
📄 Verifiable Social Reasoning for LLM Assistants
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situati…
- 2026-09-14·GitHub
What's Changed MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. Improved MLX memory handling on Apple Silicon Runa…
- 2026-09-09·GitHub
🚀 transformers v5.17.0 released
Release v5.17.0 New Model additions HYV4 <img width="1503" height="827" alt="image" src="https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9" / Hy4-Preview is a 780B-parameter mixture-of-exper…
- 2026-09-09·GitHub
v0.29.0 Highlights This release features 594 commits from 277 contributors (91 new)! Model Runner V2 is now the default for all models (53183), completing the rollout that began with pooling models (48290). MRV2 also gai…
- 2026-09-05·GitHub
Highlights 786 PRs from 214 contributors. New models in this release (see the cookbook(https://docs.sglang.io/cookbook) for all supported models): | Model | Type | PRs | Cookbook | |---|---|---|---| | Qwen3.8 (2.4T-A95B)…
- 2026-09-04·HF Papers
📄 Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
We present Srijika, a system for producing installable OpenType fonts for nine Brahmic scripts: Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Odia. Rather than generating fonts from scra…
- 2026-08-26·GitHub
🚀 transformers v5.16.1 released
Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash <img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f94…
- 2026-08-26·GitHub
🚀 transformers v5.16.0 released
Release v5.16.0 New Model additions Qwen4-Exp <img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" / Qwen4-Exp builds on Qwen3.5's hybrid text a…
- 2026-08-26·GitHub
v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)! Kimi-K3 performance push: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (50484), fus…
- 2026-08-22·GitHub
Highlights 710 PRs from 212 contributors. New models in this release (see the cookbook(https://docs.sglang.io/cookbook) for all supported models): | Model | Type | PRs | Cookbook | |---|---|---|---| | Muse Glimmer | Auto…
- 2026-08-11·GitHub
This is a patch release on top of v0.27.0. - Support quantized DSpark Markov heads (50424)