A New AI Research from Anthropic and Thinking Machines Lab Stress Tests Model Specs and Reveal Character Differences among Language Models
ตุลาคม 26, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)AI companies use model specifications to define target behaviors during training and evaluation. Do current...

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
มีนาคม 23, 2026admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2601.03273v2 Announce Type: replace Abstract: As large language models (LLMs) become deeply embedded in daily...
A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs
มิถุนายน 26, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2506.20073v1 Announce Type: new Abstract: Spatio-temporal data mining plays a pivotal role in informed decision...
A methodological analysis of prompt perturbations and their effect on attack success rates
พฤศจิกายน 17, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2511.10686v1 Announce Type: new Abstract: This work aims to investigate how different Large Language Models...
A Method for Characterizing Disease Progression from Acute Kidney Injury to Chronic Kidney Disease
พฤศจิกายน 19, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2511.14603v1 Announce Type: new Abstract: Patients with acute kidney injury (AKI) are at high risk...
A Matter of Representation: Towards Graph-Based Abstract Code Generation
ตุลาคม 16, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2510.13163v1 Announce Type: new Abstract: Most large language models (LLMs) today excel at generating raw...
A long-abandoned US nuclear technology is making a comeback in China
พฤษภาคม 1, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)China has once again beat everyone else to a clean energy milestone—its new nuclear reactor...
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
มิถุนายน 17, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2502.13520v2 Announce Type: replace Abstract: This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC)...
A Human Behavioral Baseline for Collective Governance in Software Projects
พฤศจิกายน 18, 2025admin NUAI,Committee,ข่าว,Uncategorized(0)arXiv:2510.08956v2 Announce Type: replace Abstract: We study how open source communities describe participation and control...