YouZum

Uncategorized

AI, Committee, 新闻, Uncategorized

Salesforce AI Introduces FOFPred: A Language-Driven Future Optical Flow Prediction Framework that Enables Improved Robot Control and Video Generation

Salesforce AI research team present FOFPred, a language driven future optical flow prediction framework that connects large vision language models with diffusion transformers for dense motion forecasting in control and video generation settings. FOFPred takes one or more images and a natural language instruction such as ‘moving the bottle from right to left’ and predicts 4 future optical flow frames that describe how every pixel is expected to move over time. https://arxiv.org/pdf/2601.10781 Future optical flow as a motion representation Optical flow is the apparent per pixel displacement between two frames. FOFPred focuses on future optical flow, which means predicting dense displacement fields for future frames given only current observations and text, without access to future images at inference. Future optical flow is a compact motion only representation. It removes static appearance and keeps only pixel level motion, so it is well suited as an intermediate state for robot control policies and as a conditioning signal for video diffusion models. Compared to predicting future RGB frames, it reduces the complexity of the output distribution and avoids modeling textures and high frequency details that are not required for motion planning. To plug into existing latent diffusion infrastructure, the research team encode optical flow as RGB images. They map flow magnitude and direction from polar form into HSV channels, then convert to RGB. The scaling of each channel is tuned so that consecutive flow frames are visually smooth and resemble animated graphics. A standard Flux.1 variational autoencoder then encodes and decodes these flow images. Unified VLM Diffusion backbone FOFPred uses a unified architecture that combines a frozen vision language model, a frozen VAE and a trainable diffusion transformer. The pipeline is: Qwen2.5-VL is used as the vision language encoder to jointly encode the caption and visual inputs. Flux.1 VAE encodes the input images and the training optical flow targets into latent tensors. An OmniGen style diffusion transformer, DiT, takes projected visual and textual features as conditional inputs and generates latent future flow sequences. Only the DiT and small MLP projectors are trained. The Qwen2.5-VL and Flux.1 weights stay frozen, which lets the model reuse image editing pretraining and multimodal reasoning ability from prior work. Temporal modeling is added by extending the RoPE positional encoding and attention blocks from two dimensional spatial positions to full spatio-temporal positions across input and output frame sequences. This gives full spatio-temporal attention without adding extra parameters, so the DiT can reuse OmniGen image pretraining directly. https://arxiv.org/pdf/2601.10781 Training on noisy web videos with relative optical flow The core model is trained on web scale human activity videos with paired captions. The research team uses the Something Something V2 dataset and the EgoDex egocentric manipulation dataset to obtain around 500,000 video caption pairs. Training uses an end to end flow matching objective in latent space. Future optical flow sequences are first computed offline, then encoded by the VAE and used as targets in a flow matching diffusion loss for the DiT. During training the method also applies classifier free guidance on both text and visual conditions and masks some frames and viewpoints to improve robustness. A critical contribution is the relative optical flow calculation used to build clean training targets from noisy egocentric videos. For each frame pair the method: Computes dense optical flow with an off the shelf estimator. Estimates camera motion via homography using deep features. Uses projective geometry to subtract camera motion and obtain object centric relative flow vectors. Filters frame pairs by selecting those where the top k percent flow magnitudes exceed a threshold, which focuses training on segments with meaningful motion. These steps are run offline at lower resolution for efficiency, then recomputed at original resolution for the final targets. The ablation study shows that static frame targets or raw flow without camera motion removal harm downstream performance, while disentangled relative flow targets give the best results. https://arxiv.org/pdf/2601.10781 Language driven robot manipulation The first downstream use case is robot control. FOFPred is finetuned on robot video caption data to predict future optical flow from both fixed and wrist mounted cameras. On top of FOFPred, the research team attach a diffusion policy network that takes predicted flow, text and robot state, and outputs continuous actions. This setup follows prior diffusion policy work but uses future optical flow instead of predicted RGB frames as the core representation. On the CALVIN ABCD benchmark, which evaluates long horizon zero shot chains of 5 language specified manipulation tasks, FOFPred reaches an average chain length of 4.48. VPP reaches 4.33 and DreamVLA reaches 4.44 under the same protocol. FOFPred also attains a Task 5 success rate of 78.7 percent, which is the best among reported methods. In a low data setting with 10 percent of CALVIN demonstrations, FOFPred still reaches 3.43 average length, higher than the 3.25 of VPP. On RoboTwin 2.0, a dual arm manipulation benchmark with 5 tasks that require both arms, FOFPred attains an average success rate of 68.6 percent. The VPP baseline reaches 61.8 percent under identical training settings. FOFPred improves success on every task in the subset. https://arxiv.org/pdf/2601.10781 Motion aware text to video generation The second downstream task is motion control in text to video generation. The research team build a two stage pipeline by connecting FOFPred with the Go with the Flow video diffusion model. FOFPred takes an initial frame and a language description of motion, predicts a sequence of future flow frames, and interpolates them into a dense motion field. Go with the Flow then uses this motion field and the initial frame to synthesize the final video, enforcing the described motion pattern. On the motion heavy Something Something V2 benchmark, the FOFPred along with Go with the Flow pipeline improves over the CogVideoX baseline under identical conditions. The method reaches SSIM 68.4, PSNR 22.26, LPIPS 28.5, FVD 75.39, KVD 11.38, and motion fidelity 0.662, which are consistently better than CogVideoX. Importantly, FOFPred only uses language and a single frame at inference, while several controllable video baselines require hand or object masks or

Salesforce AI Introduces FOFPred: A Language-Driven Future Optical Flow Prediction Framework that Enables Improved Robot Control and Video Generation Read Post »

AI, Committee, 新闻, Uncategorized

Going beyond pilots with composable and sovereign AI

Today marks an inflection point for enterprise AI adoption. Despite billions invested in generative AI, only 5% of integrated pilots deliver measurable business value and nearly one in two companies abandons AI initiatives before reaching production. DOWNLOAD THE ARTICLE The bottleneck is not the models themselves. What’s holding enterprises back is the surrounding infrastructure: Limited data accessibility, rigid integration, and fragile deployment pathways prevent AI initiatives from scaling beyond early LLM and RAG experiments. In response, enterprises are moving toward composable and sovereign AI architectures that lower costs, preserve data ownership, and adapt to the rapid, unpredictable evolution of AI—a shift IDC expects 75% of global businesses to make by 2027. The concept to production reality AI pilots almost always work, and that’s the problem. Proofs of concept (PoCs) are meant to validate feasibility, surface use cases, and build confidence for larger investments. But they thrive in conditions that rarely resemble the realities of production. Source: Compiled by MIT Technology Review Insights with data from Informatica, CDO Insights 2025 report, 2026 “PoCs live inside a safe bubble” observes Cristopher Kuehl, chief data officer at Continent 8 Technologies. Data is carefully curated, integrations are few, and the work is often handled by the most senior and motivated teams. The result, according to Gerry Murray, research director at IDC, is not so much pilot failure as structural mis-design: Many AI initiatives are effectively “set up for failure from the start.” Download the article.

Going beyond pilots with composable and sovereign AI Read Post »

AI, Committee, 新闻, Uncategorized

The Download: the US digital rights crackdown, and AI companionship

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. What it’s like to be banned from the US for fighting online hate   Just before Christmas the Trump administration dramatically escalated its war on digital rights by banning five people from entering the US. One of them, Josephine Ballon, is a director of HateAid, a small German nonprofit founded to support the victims of online harassment and violence. The organization is a strong advocate of EU tech regulations, and so finds itself attacked in campaigns from right-wing politicians and provocateurs who claim that it engages in censorship.  EU officials, freedom of speech experts, and the five people targeted all flatly reject these accusations. Ballon told us that their work is fundamentally about making people feel safer online. But their experiences over the past few weeks show just how politicized and besieged their work in online safety has become. Read the full story.  —Eileen Guo TR10: AI companions Chatbots are skilled at crafting sophisticated dialogue and mimicking empathetic behavior. They never get tired of chatting. It’s no wonder, then, that so many people now use them for companionship—forging friendships or even romantic relationships.  72% of US teenagers have used AI for companionship, according to a study from the nonprofit Common Sense Media. But while chatbots can provide much-needed emotional support and guidance for some people, they can exacerbate underlying problems in others—especially vulnerable people or those with mental health issues.  Although some early attempts to regulate this space are underway, AI companionship is going nowhere. Read why we made it one of our 10 Breakthrough Technologies this year, and check out the rest of the list. And, if you want to learn more about what we predict for AI this year, sign up to join me for our free LinkedIn Live event tomorrow at 12.30pm ET. Why inventing new emotions feels so good   Have you ever felt “velvetmist”?   It’s a “complex and subtle emotion that elicits feelings of comfort, serenity, and a gentle sense of floating.” It’s peaceful, but more ephemeral and intangible than contentment. It might be evoked by the sight of a sunset or a moody, low-key album.   If you haven’t ever felt this sensation—or even heard of it—that’s not surprising. A Reddit user generated it with ChatGPT, along with advice on how to evoke the feeling. Don’t scoff: Researchers say more and more terms for these “neo-­emotions” are showing up online, describing new dimensions and aspects of feeling. Read our story to learn more about why.  —Anya Kamenetz This story is from the latest print issue of MIT Technology Review. If you haven’t already, subscribe now to receive the next edition as soon as it lands (and benefit from some hefty seasonal discounts too!) The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Ads are coming to ChatGPT For American users initially, with plans to expand soon. (CNN)+ Here’s how they’ll work. (Wired $) 2 What will we be able to salvage after the AI bubble bursts? It will be ugly, but there are plenty of good uses for AI that we’ll want to keep. (The Guardian) + What even is the AI bubble? (MIT Technology Review) 3 It’s almost impossible to mine Greenland’s natural resources It has vast supplies of rare earth elements, but its harsh climate and environment make them very hard to access. (The Week) 4 Iran is now 10 days into its internet shutdownIt’s one of the longest and most extreme we’ve ever witnessed. (BBC)+  Starlink isn’t proving as helpful as hoped as the regime finds ways to jam it. (Reuters $)+ Battles are raging online about what’s really going on inside Iran. (NYT $) 5 America is heading for a polymarket disaster Prediction markets are getting out of control, and some people are losing a lot of money. (The Atlantic $)+ They were first embraced by political junkies, but now they’re everywhere. (NYT $) 6 How to fireproof a city Californians are starting to fight fires before they can even start. (The Verge $)+ How AI can help spot wildfires. (MIT Technology Review) 7 Stoking ‘deep state’ conspiracy theories can be dangerous Especially if you’re then given the task of helping run one of those state institutions, as Dan Bongino is now learning. (WP $)+ Why everything is a conspiracy now. (MIT Technology Review) 8 Why we’re suddenly all having a ‘Very Chinese Time’ It’s a fun, flippant trend—but it also shows how China’s soft power is growing around the globe. (Wired $)  9 Why there’s no one best way to store informationEach one involves trade-offs between space and time. (Quanta $) 10 Meat may play a surprising role in helping people reach 100Perhaps because it can assist with building stronger muscles and bones. (New Scientist $) Quote of the day “That’s the level of anxiety now – people watching the skies and the seas themselves because they don’t know what else to do.” —A Greenlander tells The Guardian just how seriously she and her fellow compatriots are taking Trump’s threat to invade their country.  One more thing KATHERINE LAM Inside a romance scam compound—and how people get tricked into being there Gavesh’s journey started, seemingly innocently, with a job ad on Facebook promising work he desperately needed. Instead, he found himself trafficked into a business commonly known as “pig butchering”—a form of fraud in which scammers form close relationships with targets online and extract money from them. The Chinese crime syndicates behind the scams have netted billions of dollars, and they have used violence and coercion to force their workers, many of them trafficked like Gavesh, to carry out the frauds from large compounds, several of which operate openly in the quasi-lawless borderlands of Myanmar. Big Tech may hold the key to breaking up the scam syndicates—if these companies can be persuaded or compelled to act. Read the full story. —Peter Guest & Emily Fishbein We can still have nice things A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.) + Blue Monday isn’t real (but it is an absolute banger of a track.) + Some great advice here about

The Download: the US digital rights crackdown, and AI companionship Read Post »

AI, Committee, 新闻, Uncategorized

How to Design a Fully Streaming Voice Agent with End-to-End Latency Budgets, Incremental ASR, LLM Streaming, and Real-Time TTS

In this tutorial, we build an end-to-end streaming voice agent that mirrors how modern low-latency conversational systems operate in real time. We simulate the complete pipeline, from chunked audio input and streaming speech recognition to incremental language model reasoning and streamed text-to-speech output, while explicitly tracking latency at every stage. By working with strict latency budgets and observing metrics such as time to first token and time to first audio, we focus on the practical engineering trade-offs that shape responsive voice-based user experiences. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser import time import asyncio import numpy as np from collections import deque from dataclasses import dataclass from typing import List, AsyncIterator from enum import Enum import matplotlib.pyplot as plt @dataclass class LatencyMetrics: audio_chunk_received: float = 0.0 asr_started: float = 0.0 asr_partial: float = 0.0 asr_complete: float = 0.0 llm_started: float = 0.0 llm_first_token: float = 0.0 llm_complete: float = 0.0 tts_started: float = 0.0 tts_first_chunk: float = 0.0 tts_complete: float = 0.0 def get_time_to_first_audio(self) -> float: return self.tts_first_chunk – self.asr_complete if self.tts_first_chunk and self.asr_complete else 0.0 def get_total_latency(self) -> float: return self.tts_complete – self.audio_chunk_received if self.tts_complete else 0.0 @dataclass class LatencyBudgets: asr_processing: float = 0.1 asr_finalization: float = 0.3 llm_first_token: float = 0.5 llm_token_generation: float = 0.02 tts_first_chunk: float = 0.2 tts_chunk_generation: float = 0.05 time_to_first_audio: float = 1.0 class AgentState(Enum): LISTENING = “listening” PROCESSING_SPEECH = “processing_speech” THINKING = “thinking” SPEAKING = “speaking” INTERRUPTED = “interrupted” We define the core data structures and state representations that allow us to track latency across the entire voice pipeline. We formalize timing signals for ASR, LLM, and TTS to ensure consistent measurement across all stages. We also establish a clear agent state machine that guides how the system transitions during a conversational turn. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser class AudioInputStream: def __init__(self, sample_rate: int = 16000, chunk_duration_ms: int = 100): self.sample_rate = sample_rate self.chunk_duration_ms = chunk_duration_ms self.chunk_size = int(sample_rate * chunk_duration_ms / 1000) async def stream_audio(self, text: str) -> AsyncIterator[np.ndarray]: chars_per_second = (150 * 5) / 60 duration_seconds = len(text) / chars_per_second num_chunks = int(duration_seconds * 1000 / self.chunk_duration_ms) for _ in range(num_chunks): chunk = np.random.randn(self.chunk_size).astype(np.float32) * 0.1 await asyncio.sleep(self.chunk_duration_ms / 1000) yield chunk We simulate real-time audio input by breaking speech into fixed-duration chunks that arrive asynchronously. We model realistic speaking rates and streaming behavior to mimic live microphone input. We use this stream as the foundation for testing downstream latency-sensitive components. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser class StreamingASR: def __init__(self, latency_budget: float = 0.1): self.latency_budget = latency_budget self.silence_threshold = 0.5 async def transcribe_stream( self, audio_stream: AsyncIterator[np.ndarray], ground_truth: str ) -> AsyncIterator[tuple[str, bool]]: words = ground_truth.split() words_transcribed = 0 silence_duration = 0.0 chunk_count = 0 async for chunk in audio_stream: chunk_count += 1 await asyncio.sleep(self.latency_budget) if chunk_count % 3 == 0 and words_transcribed < len(words): words_transcribed += 1 yield ” “.join(words[:words_transcribed]), False audio_power = np.mean(np.abs(chunk)) silence_duration = silence_duration + 0.1 if audio_power < 0.05 else 0.0 if silence_duration >= self.silence_threshold: await asyncio.sleep(0.2) yield ground_truth, True return yield ground_truth, True We implement a streaming ASR module that produces partial transcriptions before emitting a final result. We progressively reveal words to reflect how modern ASR systems operate in real time. We also introduce silence-based finalization to approximate end-of-utterance detection. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser class StreamingLLM: def __init__(self, time_to_first_token: float = 0.3, tokens_per_second: float = 50): self.time_to_first_token = time_to_first_token self.tokens_per_second = tokens_per_second async def generate_response(self, prompt: str) -> AsyncIterator[str]: responses = { “hello”: “Hello! How can I help you today?”, “weather”: “The weather is sunny with a temperature of 72°F.”, “time”: “The current time is 2:30 PM.”, “default”: “I understand. Let me help you with that.” } response = responses[“default”] for key in responses: if key in prompt.lower(): response = responses[key] break await asyncio.sleep(self.time_to_first_token) for word in response.split(): yield word + ” ” await asyncio.sleep(1.0 / self.tokens_per_second) class StreamingTTS: def __init__(self, time_to_first_chunk: float = 0.2, chars_per_second: float = 15): self.time_to_first_chunk = time_to_first_chunk self.chars_per_second = chars_per_second async def synthesize_stream(self, text_stream: AsyncIterator[str]) -> AsyncIterator[np.ndarray]: first_chunk = True buffer = “” async for text in text_stream: buffer += text if len(buffer) >= 20 or first_chunk: if first_chunk: await asyncio.sleep(self.time_to_first_chunk) first_chunk = False duration = len(buffer) / self.chars_per_second yield np.random.randn(int(16000 * duration)).astype(np.float32) * 0.1 buffer = “” await asyncio.sleep(duration * 0.5) In this snippet, we model a streaming language model and a streaming text-to-speech engine working together. We generate responses token by token to capture time-to-first-token behavior. We then convert incremental text into audio chunks to simulate early and continuous speech synthesis. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser class StreamingVoiceAgent: def __init__(self, latency_budgets: LatencyBudgets): self.budgets = latency_budgets self.audio_stream = AudioInputStream() self.asr = StreamingASR(latency_budgets.asr_processing) self.llm = StreamingLLM( latency_budgets.llm_first_token, 1.0 / latency_budgets.llm_token_generation ) self.tts = StreamingTTS( latency_budgets.tts_first_chunk, 1.0 / latency_budgets.tts_chunk_generation ) self.state = AgentState.LISTENING self.metrics_history: List[LatencyMetrics] = [] async def process_turn(self, user_input: str) -> LatencyMetrics: metrics = LatencyMetrics() start_time = time.time() metrics.audio_chunk_received = time.time() – start_time audio_gen = self.audio_stream.stream_audio(user_input) metrics.asr_started = time.time() – start_time async for text, final in self.asr.transcribe_stream(audio_gen, user_input): if final: metrics.asr_complete = time.time() – start_time transcription = text metrics.llm_started = time.time() – start_time response = “” async for token in self.llm.generate_response(transcription): if not metrics.llm_first_token: metrics.llm_first_token = time.time() – start_time response += token metrics.llm_complete = time.time() – start_time metrics.tts_started = time.time() – start_time async def text_stream(): for word in response.split(): yield word + ” ” async for _ in self.tts.synthesize_stream(text_stream()): if not metrics.tts_first_chunk: metrics.tts_first_chunk = time.time() – start_time metrics.tts_complete = time.time() – start_time self.metrics_history.append(metrics) return metrics We orchestrate the full voice agent by wiring audio input, ASR, LLM, and TTS into a single asynchronous flow. We record precise timestamps at each transition to compute critical latency metrics. We treat each user turn as an isolated experiment to enable systematic performance analysis. Check out the FULL CODES here. Copy CodeCopiedUse a different Browser async def run_demo(): budgets = LatencyBudgets( asr_processing=0.08, llm_first_token=0.3, llm_token_generation=0.02, tts_first_chunk=0.15, time_to_first_audio=0.8 ) agent = StreamingVoiceAgent(budgets) inputs = [ “Hello, how are you today?”, “What’s

How to Design a Fully Streaming Voice Agent with End-to-End Latency Budgets, Incremental ASR, LLM Streaming, and Real-Time TTS Read Post »

AI, Committee, 新闻, Uncategorized

Microsoft Research Releases OptiMind: A 20B Parameter Model that Turns Natural Language into Solver Ready Optimization Models

Microsoft Research has released OptiMind, an AI based system that converts natural language descriptions of complex decision problems into mathematical formulations that optimization solvers can execute. It targets a long standing bottleneck in operations research, where translating business intent into mixed integer linear programs usually needs expert modelers and days of work. What OptiMind Is And What It Outputs? OptiMind-SFT is a specialized 20B parameter Mixture of Experts model in the gpt oss transformer family. About 3.6B parameters are active per token, so inference cost is closer to a mid sized model while keeping high capacity. The context length is 128,000 tokens, which allows long specifications and multi step reasoning traces inside a single request. The model takes a natural language description of an optimization problem as input. The output is a mathematical formulation along with executable Python code that uses GurobiPy. The generated script defines decision variables, constraints, and objective, calls the Gurobi solver, and prints the optimal objective value and decisions. OptiMind acts as a formulation layer between domain experts and standard MILP solvers. It does not replace the solver, it generates the MILP that the solver will optimize. Architecture, Training Setup, And Datasets The base model is openai/gpt-oss-20b, fine tuned into microsoft/OptiMind-SFT using cleaned optimization datasets. The architecture is a Mixture of Experts transformer, with routing that activates a subset of experts per token. The model is released under the MIT license. Training uses 8 NVIDIA B200 GPUs, and inference and evaluation in the reference setup use 8 NVIDIA H100 GPUs. Reported fine tuning time is about 8 hours. For regular use, the team recommends at least 32 GB of GPU memory on hardware such as A100, H100, or B200. For supervised fine tuning, the research team construct cleaned versions of OR Instruct and OptMATH Train. For testing, they use expert validated and re-cleaned versions of IndustryOR, Mamo Complex, and OptMATH. These benchmarks cover hard formulation tasks where existing models often reach only 20 to 50 percent accuracy on the original noisy versions. Class Based Error Analysis And Data Cleaning A key technical idea in OptiMind is to combine optimization expertise with LLM training. The research team classifies problems from OR-Instruct and OptMATH into 53 seed classes, for example set cover, flow shop scheduling, or traveling salesman problem. For each class, they run the gpt-oss-20b-base model on a sample of problems and select instances where the model output disagrees with the ground truth. Optimization experts inspect these items, identify the recurring formulation mistakes, and write short error descriptions and preventive hints. These hints describe correct constraints, variable bounds, or modeling tricks, such as the proper Miller Tucker Zemlin constraints for TSP. The research team then uses a semi-automated pipeline. They regenerate solutions with a larger model that is prompted with the class specific hints, apply majority voting across samples to improve solution quality, and drop items that remain inconsistent. They also detect missing parameters and ambiguous statements and regenerate problem descriptions when needed. The result is a cleaned training corpus that is better aligned with correct mathematical formulations. Inference Pipeline, Hints, And Test Time Scaling At inference time, OptiMind behaves as a multi stage system, not just a single prompt. The default pipeline first classifies each test instance into one of the 53 optimization classes used during error analysis. It then augments the prompt with the error summary and hint pairs associated with that class. The model then generates a reasoning trace, the mathematical formulation, and the GurobiPy code. When more compute is available, the system can apply self consistency with majority voting. It generates several candidate scripts, executes them, and selects the solution that appears most often within set numerical tolerances. A multi turn correction mode can also be enabled. The system runs the generated code, captures solver logs or execution errors, feeds this feedback back to the model, and lets the model revise the formulation and code for a few rounds. This closes some modeling and coding errors at the cost of higher latency. Quantitative Gains On Optimization Benchmarks On cleaned versions of IndustryOR, Mamo-Complex, and OptMATH, the OptiMind framework significantly improves solution accuracy. The fine-tuned model improves formulation accuracy by 20.7 percent across multiple optimization benchmarks, with further gains when test time scaling techniques such as self consistency and multi turn feedback are applied. Across these benchmarks, OptiMind improves absolute accuracy over the gpt-oss-20b-base model and outperforms other open source models of similar or larger size. It reaches performance that is competitive with proprietary frontier models such as GPT-o4 mini and GPT-5 under the evaluation settings. These results rely on careful cleaning of both training and test data. The research team report that many apparent model errors on original benchmarks actually came from missing data, ambiguous descriptions, or incorrect reference solutions, and that re-cleaning can lift apparent accuracy for a fixed model from about 40 to 60 percent into the 70 to 90 percent range on the corrected sets. Key Takeaways OptiMind is a 20B parameter Mixture of Experts transformer in the gpt-oss-family that takes natural language optimization problems as input and outputs both a mathematical formulation and executable GurobiPy code, with about 3.6B parameters activated per token and a 128,000 token context length. The model is fine tuned from openai/gpt-oss-20b on cleaned optimization datasets such as OR-Instruct and OptMATH, and evaluated on expert validated benchmarks including IndustryOR and Mamo Complex, focusing on mixed integer linear programming formulations. OptiMind uses class based error analysis and expert written hints for 53 optimization classes, then applies these hints both in data cleaning and at inference time, which systematically reduces common modeling mistakes in generated MILPs. The framework improves formulation accuracy by 20.7 percent across multiple optimization benchmarks compared to the base model, and with test time scaling methods such as self consistency and multi turn feedback it reaches performance that is competitive with larger proprietary systems. OptiMind-SFT is released as microsoft/OptiMind-SFT on Hugging Face and as microsoft-optimind-sft in Azure AI Foundry, where it can be

Microsoft Research Releases OptiMind: A 20B Parameter Model that Turns Natural Language into Solver Ready Optimization Models Read Post »

AI, Committee, 新闻, Uncategorized

From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models

arXiv:2601.11020v1 Announce Type: new Abstract: Advances in mechanistic interpretability have identified special attention heads, known as retrieval heads, that are responsible for retrieving information from the context. However, the role of these retrieval heads in improving model performance remains unexplored. This work investigates whether retrieval heads can be leveraged to enhance the long-context capabilities of LLMs. Specifically, we propose RetMask, a method that generates training signals by contrasting normal model outputs with those from an ablated variant in which the retrieval heads are masked. This mechanism-based approach achieves substantial improvements: +2.28 points on HELMET at 128K for Llama-3.1, with +70% gains on generation with citation and +32% on passage re-ranking, while preserving performance on general tasks. Experiments across three model families reveal that the effectiveness depends on retrieval head organization: models with concentrated patterns of retrieval heads respond strongly, while those with distributed patterns show limited gains. This mechanistic relationship validates the function of retrieval heads and demonstrates that mechanistic insights can be transformed into performance enhancements.

From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models Read Post »

AI, Committee, 新闻, Uncategorized

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

arXiv:2512.17776v2 Announce Type: replace Abstract: As large language models advance, deep research systems capable of generating expert-level reports through multi-step reasoning and evidence-based synthesis are emerging. However, evaluating such reports remains challenging. Existing benchmarks often lack systematic evaluation criteria, rely heavily on LLM-based judges that may miss issues requiring expert judgment, and verify only a limited subset of explicitly cited statements rather than report-wide factual reliability. To address these limitations, we introduce DEER, a benchmark for evaluating expert-level deep research reports. DEER comprises 50 report-writing tasks spanning 13 domains, along with an expert-grounded evaluation taxonomy with seven dimensions and 25 subdimensions, operationalized into 101 fine-grained rubric items. To improve evaluation consistency, DEER provides task-specific Expert Evaluation Guidance to support LLM-based judging. Complementing rubric-based assessment, we propose a document-level fact-checking architecture that verifies both cited and uncited claims and quantifies the quality and reliability of the supporting evidence. Experimental results show that DEER exhibits strong correlation with human expert judgments and yields interpretable diagnostics of system strengths and weaknesses.

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation Read Post »

AI, Committee, 新闻, Uncategorized

An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit

arXiv:2601.10321v2 Announce Type: replace Abstract: Finding the most relevant person for a job proposal in real time is challenging, especially when resumes are long, structured, and multilingual. In this paper, we propose a re-ranking model based on a new generation of late cross-attention architecture, that decomposes both resumes and project briefs to efficiently handle long-context inputs with minimal computational overhead. To mitigate historical data biases, we use a generative large language model (LLM) as a teacher, generating fine-grained, semantically grounded supervision. This signal is distilled into our student model via an enriched distillation loss function. The resulting model produces skill-fit scores that enable consistent and interpretable person-job matching. Experiments on relevance, ranking, and calibration metrics demonstrate that our approach outperforms state-of-the-art baselines.

An Efficient Long-Context Ranking Architecture With Calibrated LLM Distillation: Application to Person-Job Fit Read Post »

AI, Committee, 新闻, Uncategorized

PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models

arXiv:2601.10532v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in human-centric applications, yet they often fail to provide substantive emotional support. While Reinforcement Learning (RL) has been utilized to enhance empathy of LLMs, existing reward models typically evaluate empathy from a single perspective, overlooking the inherently bidirectional interaction nature of empathy between the supporter and seeker as defined by Empathy Cycle theory. To address this limitation, we propose Psychology-grounded Empathetic Reward Modeling (PERM). PERM operationalizes empathy evaluation through a bidirectional decomposition: 1) Supporter perspective, assessing internal resonation and communicative expression; 2) Seeker perspective, evaluating emotional reception. Additionally, it incorporates a bystander perspective to monitor overall interaction quality. Extensive experiments on a widely-used emotional intelligence benchmark and an industrial daily conversation dataset demonstrate that PERM outperforms state-of-the-art baselines by over 10%. Furthermore, a blinded user study reveals a 70% preference for our approach, highlighting its efficacy in generating more empathetic responses. Our code, dataset, and models are available at https://github.com/ZhengWwwq/PERM.

PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models Read Post »

We use cookies to improve your experience and performance on our website. You can learn more at 隱私權政策 and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
zh_CN