OpenHands announced that the MiniMaxAI M2.5 model has 230 billion parameters, with 10 billion active parameters. Currently, the model is not yet available on Hugging Face. The news was shared via a Reddit post.
Google reveals that actors attempted to extract knowledge from its Gemini model via extensive prompting, aiming to train cheaper copycat models. The company defines these illicit activities as intellectual property theft, raising questions about the training data origins of the models.
Ant Group has released Ming-flash-omni-2.0, a multimodal model with 100 billion parameters (6 billion active). This unified model handles image, text, video, and audio inputs, generating outputs in the same formats. The architecture promises integrated management of various data modalities.
GPT-5.3-Codex-Spark: Our First Real-Time Coding Model Offers a 15% Speed Increase and 128k Token Context Window
OpenAI announced a new version of its Codex coding tool, highlighting it as a milestone in its relationship with a chipmaker. No details were provided on the chip's technical specifications or the performance improvements achieved.
Minimax has officially announced the release of its new language model, M2.5. Early benchmarks show promising results in several tests, including SWE-Bench and BrowseComp. The company has published a dedicated webpage with more details on the model and its capabilities. This release may be of interest to those looking for alternatives to more established models.
inclusionAI has announced the release of Ring-1T-2.5, a new large language model (LLM) designed to deliver state-of-the-art performance in tasks requiring deep thinking. The model is available on Hugging Face in FP8 format, facilitating its use and integration.
Google introduces Gemini 3 Deep Think, an update designed to navigate the complex challenges of modern science, advanced research, and precision engineering. The initiative aims to provide enhanced tools and resources for professionals in these fields.
Ovis2.6-30B-A3B, a multimodal language model (MLLM) building on Ovis2.5, has been released. This model introduces a Mixture-of-Experts (MoE) architecture to improve multimodal performance and understanding of long contexts and complex documents, while keeping management costs low.
Samsung proposes REAM (REAP-less) as an alternative to Cerebras' REAP for reducing the size of large language models (LLMs). REAM aims to minimize the loss of model capabilities during the compression process. Qwen3 models reduced via REAM have been released, opening new avenues for efficient inference. The impact of quantization and fine-tuning on REAM models remains to be evaluated.
A Reddit post expresses gratitude towards Chinese developers for their contribution to the LocalLLaMA community. The discussion highlights how their work has enabled significant progress in the field of large language models (LLMs) locally.
Z.ai has announced GLM-5, a new version of its large language model (LLM), with improvements in AI agent capabilities and a focus on compatibility with Chinese hardware. This development could have significant implications for the AI landscape in China.
A novel approach to Key-Value (KV) cache management in Large Language Models (LLMs) employs reinforcement learning (RL) to optimize token eviction. KV Policy (KVP) trains lightweight RL agents to predict the future utility of tokens, outperforming traditional heuristics and improving performance on long-context and multi-turn dialogue benchmarks.
A novel approach, Latent Thoughts Tuning (LT-Tuning), aims to enhance the reasoning capabilities of Large Language Models (LLMs) by leveraging continuous latent spaces. This method contrasts with the traditional Chain-of-Thought (CoT) approach, which constrains reasoning to the discrete space of textual vocabulary, addressing issues of feature collapse and instability.
A new mathematical research agent, Aletheia, powered by an advanced version of Gemini, is capable of generating, verifying, and revising mathematical solutions in natural language. Aletheia has demonstrated capabilities ranging from Mathematical Olympiad problems to PhD-level exercises, up to the production of scientific publications with minimal human intervention.
Researchers evaluated the ability of LLMs (BERT, NYUTron, Llama-3.1-8B, MedGemma-4B) to predict the modified Rankin Scale (mRS) after acute ischemic stroke. Fine-tuning Llama achieved promising performance, comparable to structured-data models, paving the way for text-based prognostic tools that can be integrated into clinical workflows.
LiveMedBench, a new benchmark for evaluating large language models (LLMs) in the medical field, has been introduced. This tool stands out for its continuous updating, the absence of data contamination, and an automated evaluation system based on specific criteria. The goal is to overcome the limitations of existing benchmarks, providing a more accurate measurement of LLM performance in real clinical settings.
Unsloth has announced the release of GLM-5 in GGUF format, paving the way for model inference on local hardware. The GGUF format facilitates the use of the model with tools like llama.cpp, making it accessible to a wide range of users and applications.
A Reddit post, accompanied by the hashtag #SaveLocalLLaMA, highlights the importance of supporting and developing large language models (LLMs) that can be run locally. The discussion emphasizes the need for open-source and self-hosted alternatives to proprietary cloud solutions, crucial for data sovereignty and customization.
The GLM-5 language model has achieved a score of 50 on the Intelligence Index, positioning itself as a leader among open-source models. The news was shared on Reddit, highlighting the growing interest in increasingly performant models accessible to the community.