新闻
新闻
Prompt Engineering for Agentic AI
You have probably spent time learning how to prompt AI well...
Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks
arXiv:2512.22966v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge...
Procedural Knowledge at Scale Improves Reasoning
arXiv:2604.01348v3 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve...
Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
arXiv:2509.26574v3 Announce Type: replace-cross Abstract: While large language models (LLMs) with reasoning capabilities are progressing...
Probing Neural Topology of Large Language Models
arXiv:2506.01042v3 Announce Type: replace Abstract: Probing large language models (LLMs) has yielded valuable insights into...
Probabilistic distances-based hallucination detection in LLMs with RAG
arXiv:2506.09886v2 Announce Type: replace Abstract: Detecting hallucinations in large language models (LLMs) is critical for...
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
arXiv:2506.14625v2 Announce Type: replace Abstract: Large Language Models (LLMs) have shown impressive moral reasoning abilities...
Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
arXiv:2511.21749v1 Announce Type: new Abstract: This paper introduces BRIES, a novel compound AI architecture designed...
Proactive defense against LLM Jailbreak
arXiv:2510.05052v2 Announce Type: replace-cross Abstract: The proliferation of powerful large language models (LLMs) has necessitated...
PRISM: Prompt-Refined In-Context System Modelling for Financial Retrieval
arXiv:2511.14130v1 Announce Type: cross Abstract: With the rapid progress of large language models (LLMs), financial...
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
arXiv:2505.08971v1 Announce Type: cross Abstract: In standard large vision-language models (LVLMs) pre-training, the model typically...
Prior Labs Releases TabPFN-2.5: The Latest Version of TabPFN that Unlocks Scale and Speed for Tabular Foundation Models
Tabular data is still where many important models run in production. Finance, healthcare, energy and...

