MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
9月 29, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2509.22461v1 Announce Type: cross Abstract: The ability to reason from audio, including speech, paralinguistic cues...
MCP-Universe benchmark shows GPT-5 fails more than half of real-world orchestration tasks
8月 23, 2025admin NUAI,Committee,ニュース,Uncategorized(0)A new benchmark from Salesforce research evaluates model and agentic performance on real-life enterprise tasks.Read...

MCP and the innovation paradox: Why open standards will save AI from itself
5月 11, 2025admin NUAI,Committee,ニュース,Uncategorized(0)Much like HTTP and REST standardized how web applications connect to services, MCP standardizes how...

McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
8月 8, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2507.02088v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly applied to various...
MBZUAI Researchers Introduce PAN: A General World Model For Interactable Long Horizon Simulation
11月 16, 2025admin NUAI,Committee,ニュース,Uncategorized(0)Most text to video models generate a single clip from a prompt and then stop...

MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
1月 12, 2026admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2601.05475v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong capabilities in general coding...
Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
10月 21, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2510.12997v2 Announce Type: replace-cross Abstract: Test-time scaling has enabled Large Language Models (LLMs) with remarkable...
Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science
6月 6, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2506.04410v1 Announce Type: cross Abstract: Contemporary approaches to assisted scientific discovery use language models to...
