Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
7月 24, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2507.16835v1 Announce Type: cross Abstract: Voice-based conversational AI systems increasingly rely on cascaded architectures combining...
Evaluating Rare Disease Diagnostic Performance in Symptom Checkers: A Synthetic Vignette Simulation Approach
6月 26, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2506.19750v2 Announce Type: replace Abstract: Symptom Checkers (SCs) provide users with personalized medical information. To...
Evaluating LLMs on Real-World Forecasting Against Expert Forecasters
8月 6, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2507.04562v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse...
Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
12月 23, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2511.07044v2 Announce Type: replace Abstract: Mental health disorders affect over one-fifth of adults globally, yet...
Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models
3月 23, 2026admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2603.20162v1 Announce Type: new Abstract: In contested domains, instruction-tuned language models must balance user-alignment pressures...
Evaluating Creative Short Story Generation in Humans and Large Language Models
5月 13, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2411.02316v5 Announce Type: replace Abstract: Story-writing is a fundamental aspect of human imagination, relying heavily...
Evaluating Autoformalization Robustness via Semantically Similar Paraphrasing
11月 18, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2511.12784v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently emerged as powerful tools...
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
6月 16, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2506.11111v1 Announce Type: new Abstract: Large Language Models (LLMs) have gained enormous attention in recent...
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
8月 26, 2025admin NUAI,Committee,ニュース,Uncategorized(0)arXiv:2406.13069v4 Announce Type: replace Abstract: How novel are texts generated by language models (LMs) relative...