Measuring Intent Comprehension in LLMs
3 月 13, 2026admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2506.16584v2 Announce Type: replace Abstract: People judge interactions with large language models (LLMs) as successful...
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
12 月 2, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2510.27378v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) outputs let us read a model’s step-by-step reasoning...
Measles is surging in the US. Wastewater tracking could help.
1 月 23, 2026admin NUAI,Committee,新闻,Uncategorized(0)This week marked a rather unpleasant anniversary: It’s a year since Texas reported a case...
Measles cases are rising. Other vaccine-preventable infections could be next.
2 月 20, 2026admin NUAI,Committee,新闻,Uncategorized(0)There’s a measles outbreak happening close to where I live. Since the start of this...
MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
9 月 29, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2509.22461v1 Announce Type: cross Abstract: The ability to reason from audio, including speech, paralinguistic cues...
MCP-Universe benchmark shows GPT-5 fails more than half of real-world orchestration tasks
8 月 23, 2025admin NUAI,Committee,新闻,Uncategorized(0)A new benchmark from Salesforce research evaluates model and agentic performance on real-life enterprise tasks.Read...

MCP and the innovation paradox: Why open standards will save AI from itself
5 月 11, 2025admin NUAI,Committee,新闻,Uncategorized(0)Much like HTTP and REST standardized how web applications connect to services, MCP standardizes how...

McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
8 月 8, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2507.02088v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly applied to various...
MBZUAI Researchers Introduce PAN: A General World Model For Interactable Long Horizon Simulation
11 月 16, 2025admin NUAI,Committee,新闻,Uncategorized(0)Most text to video models generate a single clip from a prompt and then stop...
