ข่าว
ข่าว
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
arXiv:2510.23558v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs), which couple acoustic perception with...
Is vibe coding ruining a generation of engineers?
AI tools are revolutionizing software development by automating repetitive tasks, refactoring bloated code, and identifying...
Is There a Community Edition of Palantir? Meet OpenPlanter: An Open Source Recursive AI Agent for Your Micro Surveillance Use Cases
The balance of power in the digital age is shifting. While governments and large corporations...
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
arXiv:2504.01928v2 Announce Type: replace Abstract: Despite their impressive capabilities, LLMs exhibit a basic generalization failure...
Is the Pentagon allowed to surveil Americans with AI?
The ongoing public feud between the Department of Defense and the AI company Anthropic has...
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
arXiv:2502.13962v2 Announce Type: replace Abstract: Scaling the test-time compute of large language models has demonstrated...
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
arXiv:2510.01367v3 Announce Type: replace-cross Abstract: Reward hacking, where a reasoning model exploits loopholes in a...
Is In-Context Learning Learning?
arXiv:2509.10414v2 Announce Type: replace Abstract: In-context learning (ICL) allows some autoregressive models to solve tasks...
Is fake grass a bad idea? The AstroTurf wars are far from over.
A rare warm spell in January melted enough snow to uncover Cornell University’s newest athletic...
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health
arXiv:2603.09416v1 Announce Type: new Abstract: Large Language Models (LLMs) excel in Natural Language Processing (NLP)...
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
arXiv:2411.15993v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities in text...
Introspective Growth: Automatically Advancing LLM Expertise in Technology Judgment
arXiv:2505.12452v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly demonstrate signs of conceptual understanding...

