YouZum

Notizie

Notizie

Meet ‘AutoAgent’: The Open-Source Library That Lets an AI Engineer and Optimize Its Own Agent Harness Overnight

There’s a particular kind of tedium that every AI engineer knows intimately: the prompt-tuning loop...

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

arXiv:2505.24040v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable performance on various...

MedKGI: Iterative Differential Diagnosis with Medical Knowledge Graphs and Information-Guided Inquiring

arXiv:2512.24181v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have demonstrated significant...

Measuring Reasoning Utility in LLMs via Conditional Entropy Reduction

arXiv:2508.20395v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) often rely on...

Measuring Intent Comprehension in LLMs

arXiv:2506.16584v2 Announce Type: replace Abstract: People judge interactions with large language models (LLMs) as successful...

Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

arXiv:2510.27378v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) outputs let us read a model’s step-by-step reasoning...

Measles is surging in the US. Wastewater tracking could help.

This week marked a rather unpleasant anniversary: It’s a year since Texas reported a case...

Measles cases are rising. Other vaccine-preventable infections could be next.

There’s a measles outbreak happening close to where I live. Since the start of this...

MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark

arXiv:2509.22461v1 Announce Type: cross Abstract: The ability to reason from audio, including speech, paralinguistic cues...

MCP-Universe benchmark shows GPT-5 fails more than half of real-world orchestration tasks

A new benchmark from Salesforce research evaluates model and agentic performance on real-life enterprise tasks.Read...

MCP and the innovation paradox: Why open standards will save AI from itself

Much like HTTP and REST standardized how web applications connect to services, MCP standardizes how...

McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

arXiv:2507.02088v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly applied to various...

We use cookies to improve your experience and performance on our website. You can learn more at Politica sulla privacy and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
it_IT