Topic / Trend Stable

AI Research Benchmarks and Specialized Evaluation

New studies test LLMs on long-context summarization, retrieval with BM25, digital twins, epidemic forecasting, neural decoding, scientific software recognition and self-directed research.

Detected: 2026-08-25 · Updated: 2026-08-25

Related Coverage

2026-08-25 ArXiv cs.CL

Khmer Semantic Search Still Rewards the Old BM25

A study on 3,000 Khmer documents shows BM25 outperforms dense retrieval and Qwen2.5-assisted query expansion. The larger model produces better expansions but still introduces noise. A signal for teams building low-cost local search.

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-24 ArXiv cs.CL

LLM Digital Twins: The Real Bottleneck Is Structure, Not Data Volume

Research shows that organizing personal information into structured schemas improves predictive accuracy for LLM-based digital twins. On homogeneous benchmarks the gain is +1.91 percentage points, but fixed schemas fail on heterogeneous tasks. An aut...

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-23 LocalLLaMA

Fine-tuning a 450M VLM: from 1 to 44 out of 100 with browser screenshots

A 450M parameter VLM goes from 1 to 44 out of 100 after fine-tuning on 50,000 browser screenshots. The source does not specify the metric or infrastructure, but the jump shows compact models can become usable on domain-specific visual data. For those...

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-21 ArXiv cs.CL

SNAIL: Bioinformatic Software Recognition Beats General-Purpose LLMs

A hybrid framework combines lexical signals and SciBERT semantics to identify software and database names in biomedical literature. Trained with a pipeline mixing citation extraction and LLM-assisted distillation, it outperforms specialized methods a...

#Hardware #Fine-Tuning
2026-08-20 Microsoft Research

Skala 1.1 expands DFT code access and introduces a living benchmark

Microsoft Research has released Skala 1.1, a deep-learning exchange-correlation functional. Trained on 2.5 times more data, it lowers the weighted average error to 2.8 kcal/mol on GMTKN55 while retaining meta-GGA cost. Native integration in CP2K, wit...

#Hardware #LLM On-Premise #DevOps
2026-08-20 ArXiv cs.CL

LongNovel tests hallucinations in summaries of novels from 16k to 100k tokens

LongNovel is a multi-scale, bilingual Chinese-English benchmark for detecting hallucinations in long novel summaries. Built on 29 Chinese novels from 16k to 100k tokens and BookSum data, it identifies eight hallucination types. The test set was manua...

#Hardware #LLM On-Premise #Fine-Tuning
2026-08-19 ArXiv cs.CL

MD-SigLIP: Semantic Alignment for Retrieval-Based Brain-Language Decoding

A new framework aligns brain and text embeddings in a shared semantic space for retrieval-based decoding, separating neural signal from LLM reconstruction. MD-SigLIP uses duplicate-aware contrastive learning and a listwise margin term to enforce rank...

#LLM On-Premise #DevOps
2026-08-18 MIT Technology Review

Self-improving AI hits a wall: agents fail open-ended research

A Princeton-led study put Claude Opus 4.8 agents to work on unpublished NeurIPS 2026 research questions. The agents handled engineering tasks but lacked the judgment and creativity needed for open-ended research, and both papers were rejected. The fi...

#Hardware #LLM On-Premise #Fine-Tuning
← Back to All Topics