ข่าว
ข่าว
Salesforce AI Introduces CRMArena-Pro: The First Multi-Turn and Enterprise-Grade Benchmark for LLM Agents
AI agents powered by LLMs show great promise for handling complex business tasks, especially in...
Sakana AI’s Error Diffusion Trains Dale-Compliant Dual-Stream Networks, Reaching 96.7% MNIST and 61.7% CIFAR-10 Without Backpropagation
Backpropagation dominates deep learning, yet it uses a mechanism the brain likely cannot. Specifically, the...
Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM
Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu...
Sakana AI Released ShinkaEvolve: An Open-Source Framework that Evolves Programs for Scientific Discovery with Unprecedented Sample-Efficiency
Table of contents What problem is it actually solving? Does the sample-efficiency claim hold beyond...
Sakana AI Launches Sakana Translate, a Namazu-Powered Japanese–English–Chinese Translation Tool With Translate, Proofread, and Ask Modes
Sakana AI has added a new feature called Sakana Translate to its chat service, Sakana...
Sakana AI Introduces Text-to-LoRA (T2L): A Hypernetwork that Generates Task-Specific LLM Adapters (LoRAs) based on a Text Description of the Task
Transformer models have significantly influenced how AI systems approach tasks in natural language understanding, translation...
Sakana AI Introduces KAME: A Tandem Speech-to-Speech Architecture That Injects LLM Knowledge in Real Time
The fundamental tension in conversational AI has always been a binary choice: respond fast or...
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
arXiv:2511.02280v1 Announce Type: cross Abstract: We introduce SAIL-RL, a reinforcement learning (RL) post-training framework that...
Safety and accuracy follow different scaling laws in clinical large language models
arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context...
SafeMT: Multi-turn Safety for Multimodal Language Models
arXiv:2510.12133v1 Announce Type: new Abstract: With the widespread use of multi-modal Large Language models (MLLMs)...
Safely Deploying ML Models to Production: Four Controlled Strategies (A/B, Canary, Interleaved, Shadow Testing)
Deploying a new machine learning model to production is one of the most critical stages...
SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
arXiv:2508.08211v1 Announce Type: new Abstract: Watermarking LLM-generated text is critical for content attribution and misinformation...





