EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
7 月 14, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2503.08893v2 Announce Type: replace Abstract: An ideal model evaluation should achieve two goals: identifying where...
EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
2 月 17, 2026admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2505.12185v5 Announce Type: replace-cross Abstract: Evaluating the programming robustness of large language models (LLMs) is...
Europe’s extreme heat is shutting down power plants
6 月 24, 2026admin NUAI,Committee,新闻,Uncategorized(0)Europe is in the middle of a record-breaking heat wave, and the grid is being...
Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
11 月 12, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2511.08109v1 Announce Type: new Abstract: This paper examines how science fiction destabilises ontological categories by...
Estimating Privacy Leakage of Augmented Contextual Knowledge in Language Models
12 月 17, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2410.03026v3 Announce Type: replace Abstract: Language models (LMs) rely on their parametric knowledge augmented with...
Estimating LLM Uncertainty with Logits
5 月 8, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2502.00290v4 Announce Type: replace Abstract: Over the past few years, Large Language Models (LLMs) have...
Establishing AI and data sovereignty in the age of autonomous systems
5 月 14, 2026admin NUAI,Committee,新闻,Uncategorized(0)When generative AI first moved from research labs into real-world business applications, enterprises made a...

Erasing Conceptual Knowledge from Language Models
7 月 23, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2410.02760v3 Announce Type: replace Abstract: In this work, we introduce Erasure of Language Memory (ELM)...
Epistemic Diversity and Knowledge Collapse in Large Language Models
10 月 31, 2025admin NUAI,Committee,新闻,Uncategorized(0)arXiv:2510.04226v4 Announce Type: replace Abstract: Large language models (LLMs) tend to generate lexically, semantically, and...