Updates
New and trending models on Hugging Face, the latest LLM research from arXiv and HF Daily Papers, and releases of the major open inference stacks on GitHub.
Last refreshed 2026-09-20 00:01 UTC · refreshes automatically every 15 minutes
- 2026-09-17·arXiv
🧪 Paint-Anything: Unified Any-Color Control for Image Generation and Editing
Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorizat…
- 2026-09-17·arXiv
🧪 Score Centering Stabilizes Off-policy Reinforcement Learning
Reinforcement learning (RL) of large language models is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely e…
- 2026-09-17·arXiv
🧪 An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effe…
- 2026-09-17·arXiv
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit d…
- 2026-09-17·arXiv
🧪 Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adop…
- 2026-09-17·arXiv
This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable …
- 2026-09-17·arXiv
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, …
- 2026-09-17·arXiv
🧪 Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-st…
- 2026-09-17·arXiv
🧪 WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution
Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also…
- 2026-09-17·arXiv
🧪 SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment
Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional…
- 2026-09-17·arXiv
Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic …
- 2026-09-17·arXiv
🧪 Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language …
- 2026-09-17·arXiv
🧪 An Analysis of Training-Free Self-Reported Confidence in Language Models
Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbaliz…
- 2026-09-17·arXiv
🧪 Parallelism, critical windows, and separations among diffusion language models
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward p…
- 2026-09-17·arXiv
🧪 Edustories: A Collection of Real-world Case Studies from Classroom Practices
Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective cl…