A New AI Research from Anthropic and Thinking Machines Lab Stress Tests Model Specs and Reveal Character Differences among Language Models
Oktober 26, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)AI companies use model specifications to define target behaviors during training and evaluation. Do current...

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
März 23, 2026admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2601.03273v2 Announce Type: replace Abstract: As large language models (LLMs) become deeply embedded in daily...
A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs
Juni 26, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2506.20073v1 Announce Type: new Abstract: Spatio-temporal data mining plays a pivotal role in informed decision...
A methodological analysis of prompt perturbations and their effect on attack success rates
November 17, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2511.10686v1 Announce Type: new Abstract: This work aims to investigate how different Large Language Models...
A Method for Characterizing Disease Progression from Acute Kidney Injury to Chronic Kidney Disease
November 19, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2511.14603v1 Announce Type: new Abstract: Patients with acute kidney injury (AKI) are at high risk...
A Matter of Representation: Towards Graph-Based Abstract Code Generation
Oktober 16, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2510.13163v1 Announce Type: new Abstract: Most large language models (LLMs) today excel at generating raw...
A long-abandoned US nuclear technology is making a comeback in China
Mai 1, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)China has once again beat everyone else to a clean energy milestone—its new nuclear reactor...
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
Juni 17, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2502.13520v2 Announce Type: replace Abstract: This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC)...
A Human Behavioral Baseline for Collective Governance in Software Projects
November 18, 2025admin NUAI,Committee,Nachrichten,Uncategorized(0)arXiv:2510.08956v2 Announce Type: replace Abstract: We study how open source communities describe participation and control...