Actualités
Actualités
Meet ‘AutoAgent’: The Open-Source Library That Lets an AI Engineer and Optimize Its Own Agent Harness Overnight
There’s a particular kind of tedium that every AI engineer knows intimately: the prompt-tuning loop...
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
arXiv:2505.24040v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable performance on various...
MedKGI: Iterative Differential Diagnosis with Medical Knowledge Graphs and Information-Guided Inquiring
arXiv:2512.24181v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have demonstrated significant...
Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction
arXiv:2508.20395v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) often rely on...
Measuring Intent Comprehension in LLMs
arXiv:2506.16584v2 Announce Type: replace Abstract: People judge interactions with large language models (LLMs) as successful...
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
arXiv:2510.27378v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) outputs let us read a model’s step-by-step reasoning...
Measles is surging in the US. Wastewater tracking could help.
This week marked a rather unpleasant anniversary: It’s a year since Texas reported a case...
Measles cases are rising. Other vaccine-preventable infections could be next.
There’s a measles outbreak happening close to where I live. Since the start of this...
MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
arXiv:2509.22461v1 Announce Type: cross Abstract: The ability to reason from audio, including speech, paralinguistic cues...
MCP-Universe benchmark shows GPT-5 fails more than half of real-world orchestration tasks
A new benchmark from Salesforce research evaluates model and agentic performance on real-life enterprise tasks.Read...
MCP and the innovation paradox: Why open standards will save AI from itself
Much like HTTP and REST standardized how web applications connect to services, MCP standardizes how...
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
arXiv:2507.02088v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly applied to various...


