It’s pretty easy to get DeepSeek to talk dirty
6 月 20, 2025admin NUAI,Committee,新闻,Uncategorized(0)AI companions like Replika are designed to engage in intimate exchanges, but people use general-purpose...
It’s the same but not the same: Do LLMs distinguish Spanish varieties?
4 月 30, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2504.20049v1 Announce Type: new Abstract: In recent years, large language models (LLMs) have demonstrated a...
It Takes Two: Your GRPO Is Secretly DPO
10 月 2, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2510.00977v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) is a prominent reinforcement learning...
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
10 月 28, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2510.23558v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs), which couple acoustic perception with...
Is vibe coding ruining a generation of engineers?
10 月 12, 2025admin NUAI,Committee,新闻,Uncategorized(0)AI tools are revolutionizing software development by automating repetitive tasks, refactoring bloated code, and identifying...

Is There a Community Edition of Palantir? Meet OpenPlanter: An Open Source Recursive AI Agent for Your Micro Surveillance Use Cases
2 月 22, 2026admin NUAI,Committee,新闻,Uncategorized(0)The balance of power in the digital age is shifting. While governments and large corporations...
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
2 月 11, 2026admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2504.01928v2 Announce Type: replace Abstract: Despite their impressive capabilities, LLMs exhibit a basic generalization failure...
Is the Pentagon allowed to surveil Americans with AI?
3 月 7, 2026admin NUAI,Committee,新闻,Uncategorized(0)The ongoing public feud between the Department of Defense and the AI company Anthropic has...
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
7 月 21, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2502.13962v2 Announce Type: replace Abstract: Scaling the test-time compute of large language models has demonstrated...