EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
7月 14, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2503.08893v2 Announce Type: replace Abstract: An ideal model evaluation should achieve two goals: identifying where...
EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
2月 17, 2026admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2505.12185v5 Announce Type: replace-cross Abstract: Evaluating the programming robustness of large language models (LLMs) is...
Europe’s extreme heat is shutting down power plants
6月 24, 2026admin NUAI,Committee,ニュース,Uncategorized(0)Europe is in the middle of a record-breaking heat wave, and the grid is being...
Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
11月 12, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2511.08109v1 Announce Type: new Abstract: This paper examines how science fiction destabilises ontological categories by...
Estimating Privacy Leakage of Augmented Contextual Knowledge in Language Models
12月 17, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2410.03026v3 Announce Type: replace Abstract: Language models (LMs) rely on their parametric knowledge augmented with...
Estimating LLM Uncertainty with Logits
5月 8, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2502.00290v4 Announce Type: replace Abstract: Over the past few years, Large Language Models (LLMs) have...
Establishing AI and data sovereignty in the age of autonomous systems
5月 14, 2026admin NUAI,Committee,ニュース,Uncategorized(0)When generative AI first moved from research labs into real-world business applications, enterprises made a...

Erasing Conceptual Knowledge from Language Models
7月 23, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2410.02760v3 Announce Type: replace Abstract: In this work, we introduce Erasure of Language Memory (ELM)...
Epistemic Diversity and Knowledge Collapse in Large Language Models
10月 31, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2510.04226v4 Announce Type: replace Abstract: Large language models (LLMs) tend to generate lexically, semantically, and...