Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models
March 23, 2026admin NUAI,Committee,News,Uncategorized(0)arXiv:2603.20162v1 Announce Type: new Abstract: In contested domains, instruction-tuned language models must balance user-alignment pressures...
Evaluating Creative Short Story Generation in Humans and Large Language Models
May 13, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2411.02316v5 Announce Type: replace Abstract: Story-writing is a fundamental aspect of human imagination, relying heavily...
Evaluating Autoformalization Robustness via Semantically Similar Paraphrasing
November 18, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2511.12784v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently emerged as powerful tools...
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
June 16, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2506.11111v1 Announce Type: new Abstract: Large Language Models (LLMs) have gained enormous attention in recent...
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
August 26, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2406.13069v4 Announce Type: replace Abstract: How novel are texts generated by language models (LMs) relative...
EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
July 14, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2503.08893v2 Announce Type: replace Abstract: An ideal model evaluation should achieve two goals: identifying where...
EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
February 17, 2026admin NUAI,Committee,News,Uncategorized(0)arXiv:2505.12185v5 Announce Type: replace-cross Abstract: Evaluating the programming robustness of large language models (LLMs) is...
Europe’s extreme heat is shutting down power plants
June 24, 2026admin NUAI,Committee,News,Uncategorized(0)Europe is in the middle of a record-breaking heat wave, and the grid is being...
Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
November 12, 2025admin NUAI,Committee,News,Uncategorized(0)arXiv:2511.08109v1 Announce Type: new Abstract: This paper examines how science fiction destabilises ontological categories by...