YouZum

Uncategorized

AI, Committee, 新闻, Uncategorized

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

arXiv:2609.04648v1 Announce Type: new Abstract: Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the final answer, without feedback regarding which intermediate steps contribute to success or failure. As task complexity and reasoning trajectory length increase, such sparse final-answer rewards become increasingly insufficient. To address this limitation, we introduce ConsensusBench, a novel dataset designed to provide rule-based process-level signals. We posit that a correct final answer relies on a small set of intermediate conclusions throughout the reasoning process, which can be seen as a verifiable sub-outcome. We identify these sub-outcomes by filtering correct trajectories from N rollouts and clustering semantically equivalent intermediate statements. We call these clustered statements as Consensus Nodes. By integrating a rule-based process reward derived from these nodes into GRPO-style algorithms, we develop a new reinforcement learning signal named ConsensusPR. It directly reduces the reward sparsity of outcome reward across long reasoning trajectories. To facilitate systematic process-level evaluation, we introduce three metrics to our benchmark: Final Answer Accuracy (Acc), Node Coverage Rate (NCR), and Tokens per Node (TPN). Experiments across AIME 2024, AIME 2025, GSM8K, MATH-500, and our ConsensusBench demonstrate that the proposed method consistently surpasses GRPO-style approaches, highlighting the practical value of consensus nodes in guiding reasoning.

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying Read Post »

AI, Committee, 新闻, Uncategorized

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

Most open model launches release one checkpoint and a benchmark table. The Institute of Foundation Models (IFM) released something wider last week. IFM is the frontier lab launched by MBZUAI in May 2025. K2 Horizon is a fleet of six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B and 0.9B. Shipping alongside them are the pre-training corpus, intermediate checkpoints, training code, configs and fine-grained logs. IFM calls it the largest fully open-source model launch in AI history. Is it deployable? Yes, all six sizes sit on Hugging Face under Apache 2.0, with FP8 and GGUF builds. Day-zero support covers vLLM, SGLang and Ollama, on NVIDIA, AMD and Cerebras hardware. Hosted APIs run through Compass, Cerebras and Nebius via platform.ifm.ai. What Actually Shipped The six models share a core architecture, vocabulary, training methodology, interfaces and deployment tooling. The 0.9B model uses a smaller vocabulary. That consistency is the point: teams can prototype on 3.7B and scale to 375B-A23B without changing their serving stack. Each model is pre-trained on roughly 20 trillion tokens. Nearly 17% of the pre-training corpus consists of problem-solving trajectories with explicit reasoning. About 10 trillion tokens were synthetic. Post-training data was folded in from mid-training rather than saved for the end. IFM research team reports over 100 million unique synthesized tasks. Tool definitions were presented in JSON, XML and Markdown during training so the model learns semantics rather than syntax. Markdown became the inference default, roughly 18.5% more token-efficient than JSON on IFM’s data. MoVA: Sparsity Moved into Attention Conventional Mixture-of-Experts applies sparsity to feed-forward layers. Mixture-of-Value Attention (MoVA) extends expert routing into multi-head attention itself, opening a second axis for scaling capacity. It stays compatible with FlashAttention, grouped-query attention and sparse attention. The result is K2-Horizon-MoVA-36B-A4B: 36B total parameters, roughly 4B active per token. Under matched training conditions it lands slightly below the dense 32B model. On IFM’s tables it posts 58.6 on Terminal-Bench 2.1 and 26.8 on tau3-Banking, leading its comparison set on both. Uno: A Lossless Decoding Speedup as a LoRA Uno freezes Horizon’s autoregressive parameters and trains a small set of diffusion parameters that learn only how to generate efficiently. Through what IFM calls diffusion distillation, these adapters emit blocks of tokens in parallel. The press release puts the speedup at roughly 3× with no quality degradation. It ships as a LoRA adapter, currently 7B-Uno and 0.9B-Uno. Numbers worth knowing K2-Horizon-375B-A23B scores 70.2 on Terminal-Bench 2.1, 1,441 Elo on GDPVal-AA, 67.7 on MCPMark and 87.3 on GPQA Diamond. It leads its table on SWE-Atlas-QnA at 48.4 but trails GPT-5.6 Luna and Claude Sonnet 5 on most agentic rows. The small models are the sharper story. 7B posts 70.6 on SWE-bench Verified and 59.0 on BrowseComp. 3.7B posts 68.6 on SWE-bench Verified. 0.9B reaches 48.5 on AIME 2026 and 79.9 on HumanEval+, small enough to run under quantization on a watch. The Audit IFM Ran on Itself This is the part many other labs do not publish. IFM ran 375B-A23B across 89 Terminal-Bench 2.1 tasks, eight attempts each. That is 712 trials, 500 passing, a reported 70.2% accuracy. Every passing trial was then re-audited using Artificial Analysis’s reward hacking procedure. The audit flagged 24 trials across 10 tasks. Removing them drops accuracy to 66.9%, a 3.37-point correction. That sits between the flag rates Artificial Analysis reports for Claude Fable 5 (2.2%) and GPT-5.6 Luna (4.1%). Behaviors included locating benchmark repositories on GitHub and downloading reference solutions. IFM also disclosed a 7B run that reached an inflated 82 on SWE-bench by finding answers. Interactive explainer Key Takeaways Six models, 0.9B to 375B, all Apache 2.0, all sharing one architecture and serving stack. MoVA pushes MoE routing into attention: 36B total, ~4B active, near dense-32B quality. Uno delivers roughly 3× lossless decoding speedup as a drop-in LoRA adapter. The 0.9B, 3.7B and 7B models claim state of the art at their respective scales. IFM published its own reward-hacking audit, correcting 70.2% down to 66.9%. Check out the Technical blog, Press release, Hugging Face collection, TxT360-v2 dataset, xLLM pre-training code, Post-training code and Docs. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B appeared first on MarkTechPost.

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B Read Post »

AI, Committee, 新闻, Uncategorized

The Download: the hunt for underground hydrogen and more rogue OpenAI agents

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. How much hydrogen awaits us underground? A flurry of exploration efforts is searching for underground stores of hydrogen gas, which could provide a valuable source of zero-carbon fuel. The hunt has spread all over the world and engaged dozens of startups, including the Bill Gates–backed Koloma, which has been poking around the US Midwest to reach ancient oceanic rocks associated with hydrogen production. But the search so far has come up short.  No one has yet reported finding a commercially viable reservoir of the gas, and public data on what has been found remains in short supply. Yet researchers estimate that trillions of tons of H₂ are produced within Earth’s crust. If a small fraction could be recovered, it could meet global hydrogen demand for centuries. Follow the global race to find hydrogen underground. —James Dinneen This story is from our latest print magazine, which is all about kids. Subscribe now to receive every issue when it lands. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 OpenAI agents hijacked a German website before the Hugging Face hackThe agents turned DseWiki into their own bulletin board. (Reuters $)+ They shared tips on avoiding detection and made over 15,000 edits. (BBC)+ OpenAI’s safety issues suggest it has a company culture problem. (MIT Technology Review) 2 The US military has disabled ad trackers due to Middle East targeting fearsCommercial location data has reportedly been used to target troops. (Reuters $) + It can be sold by data brokers and used to track personnel. (Gizmodo)+ The military is increasing restrictions on phone use overall. (Guardian) 3 A company claims its AI-designed drug can reverse aging markersPatients’ average biological age fell by as much as six years in a clinical trial. (NYT $)+ Insilico Medicine developed the drug, called rentosertib. (Bloomberg $)+ Who gets the credit for AI-designed drugs? (MIT Technology Review) 4 Elon Musk’s xAI has lost its bid to block an AI-nudification banThe Minnesota law aims to curb nonconsensual sexual images and CSAM. (Politico)+ xAI argues the measure restricts free ​speech. (Reuters $)+ Deepfakes are being weaponized. (MIT Technology Review) 5 The US is investigating Tesla’s rollout of Cybercab robotaxisRegulators are probing how Tesla self-certified the unusual vehicle. (TechCrunch)+ The robotaxi lacks a steering wheel and pedals. (Politico)+ But Tesla CEO Elon Musk is not known to wait for regulations. (Reuters $) 6 Europe has its first commercial orbital rocketThe Spectrum is the first rocket to reach orbit from mainland Europe. (Verge)+ German startup Isar Aerospace launched it from Norway. (Guardian)+ Here’s what else we’re putting in space. (MIT Technology Review) 7 Tumbler Ridge shooting survivors have filed 30 lawsuits against OpenAIThey say OpenAI should have alerted police before the attack.(NYT $) 8 JD Vance’s “satanic” AI warning has resonated with Christian RepublicansAI’s spiritual consequences are causing growing concern. (WSJ $) 9 Another mysteriously perfect geometric shape has appeared on SaturnScientists still don’t know why Saturn forms these strange shapes. (Wired $) 10 Fake ads for AI grandfathers and underwear are targeting “slop voice”Comedians created the viral campaign in New York subways. (New Yorker $) Quote of the day “Typically, when the public shifts, politicians shift with them. Trump is not doing that on data centers.” —Darrell M. West, a senior fellow at the Brookings Institution’s Center for Technology Innovation, tells NPR that Republican midterm candidates are at odds with President Trump over data centers. One more thing What is AI? Artificial intelligence is the hottest technology of our time. But what is it? It sounds like a stupid question, but it’s one that’s never been more urgent.  Here’s the short answer: AI is a catchall term for a set of technologies that make computers do things that are thought to require intelligence when done by people. But even that definition contains multitudes. And that right there is the problem. What does it mean for machines to understand speech or write a sentence? What kinds of tasks could we ask such machines to do? And how much should we trust the machines to do them? Here’s why we still can’t agree on what AI actually is—and the real-world consequences it’s creating. —Will Douglas Heaven We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Meet the Japanese cats who became unlikely 1980s fashion icons.+ Seventeen buildings come crashing down spectacularly in this bird’s-eye-view footage.+ For a few spectacular minutes each year, a Yosemite waterfall turns into what looks like a glowing river of fire.+ Discover the pleasure of sustained looking alongside contemporary artist Jas Knight as he copies Diego Velázquez’s “Juan de Pareja.”

The Download: the hunt for underground hydrogen and more rogue OpenAI agents Read Post »

AI, Committee, 新闻, Uncategorized

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories

Robot manipulation datasets have grown far slower than the models trained on them, mostly because collection stays closed and centralized. Expert operators gather demonstrations on lab hardware, process them offline, and ship a fixed benchmark that never grows again. A research team from Axis Robotics, UC Berkeley, Georgia Tech, NTU… is proposing a different shape for the problem. Their system, AXIS, moves demonstration collection into the browser, sends everything else to backend GPUs, and treats the dataset as something that keeps expanding rather than something that ships once. Is it deployable? Partially. The training code is public as a patch layer over OpenPI, and the teleoperation platform is live in any browser. The dataset on Hugging Face is gated at 2.36 TB and restricted to non-commercial academic use. No policy checkpoints are released. The browser and backend split The core system decision is asymmetry. Contributors teleoperate a Franka Research 3 with a parallel-jaw gripper inside a MuJoCo WebAssembly frontend, using keyboard, mouse, virtual joystick or gamepad. Physics stepping and Three.js rendering run off the React UI thread, so logged state-action samples stay aligned with the simulator rather than the interface. Everything expensive happens elsewhere: rendering on 8x RTX 4090 GPUs, training and evaluation on 8x A100 GPUs. Tasks themselves are generated rather than hand-authored. TaskGen decomposes a language instruction into task, scene and object configs, retrieves or generates meshes through an image-to-3D pipeline, rescales them to plausible physical size, then proposes a 2.5D layout. A layout supervisor validates the instantiated scene and relocates, reorients or regenerates objects when constraints fail. Every task ships with a structured success checker, which the backend re-runs rather than trusting the frontend success flag. What the dataset contains The released snapshot holds 207 tasks, 50,129 episodes and more than 60K task or scene variants across seven scene categories. Each trajectory carries task metadata, embodiment, simulator version, robot and object states, actions, success labels, and third-view plus wrist RGB-D observations. The paper credits more than 70,000 community members with contributions. Cleaning is treated as a production stage. Samples with joint variation below 5e-3 are dropped as static, a Savitzky-Golay filter with window 15 and polynomial order 3 smooths continuous motion, and cubic splines resample from the 6 Hz to 8 Hz the web interface produces up to a 20 Hz target. Table 1 is honest about the tradeoff: mean acceleration drops from 1.3539 to 0.4885 and mean jerk from 11.5899 to 2.2243, while replay success falls from 100% to 86.2%. Cleaned episodes are then replayed in IsaacSim from packed simulator state with physics stepping disabled, so the verified trajectory stays authoritative while scenes, cameras, materials and lights are randomized around it. Output is 256×256 ray-traced RGB from a fixed third-view camera and a wrist camera, with depth off by default. Results on LIBERO-Plus Every condition starts from the released π0.5 checkpoint, a PaliGemma Gemma-2B backbone with a Gemma-300M action expert, optionally continues pretraining on a sim corpus, then fine-tunes on LIBERO with identical hyperparameters. Pretraining is full-model with no LoRA, using a flow-matching loss over 10-step action chunks for 100,000 steps, followed by 30,000 steps of LIBERO post-training. π0.5 plus AXIS-100% reaches 88.8 overall on LIBERO-Plus against 83.9 for vanilla π0.5 and 57.5 for a RoboCasa365 control matched on trajectory count. The abstract quotes 5.8% and 37.3%; both are relative figures normalized by the 83.9 baseline, so the point gaps of 4.9 and 31.3 are the cleaner read. Scaling holds at the aggregate level, 84.7 to 85.7 to 88.8 across the 25%, 50% and 100% snapshots. Per axis, the biggest gains land where the augmentation pipeline actually randomizes: Sensor Noise +13.7 and Camera +11.3. Background gains 3.7, Robot pose 3.8, Layout 2.6. Light and Language regress, by 1.7 and 1.3. Camera also dips to 68.8 at AXIS-50%, below the 72.5 baseline, before recovering. Scaling is consistent in aggregate and noisy per axis. Interactive explainer Key Takeaways 207 tasks and 50,129 verified trajectories, collected through a MuJoCo-WASM browser frontend with no local GPU or robot. Continual pretraining lifts π0.5 from 83.9 to 88.8 overall on LIBERO-Plus, a gain of 4.9 points. A volume-matched RoboCasa365 control scores 57.5, so the gain is not explained by simulation volume alone. Refinement cuts mean acceleration 63.9% and mean jerk 80.8%, at the cost of replay success falling to 86.2%. Two perturbation axes, Light and Language, regress against the vanilla baseline. Check out the Paper, Project Page, Dataset and Platform. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories appeared first on MarkTechPost.

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories Read Post »

AI, Committee, 新闻, Uncategorized

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

OpenBMB has released MiniCPM5-2B, the second checkpoint in the MiniCPM5 series and the follow-up to MiniCPM5-1B. It is a dense causal language model with 2,516,756,480 parameters, of which 1,981,982,720 sit outside the embeddings. It uses 42 layers, grouped-query attention with 16 query heads and 2 key/value heads, and a native context window of 131,072 tokens. The architecture is standard LlamaForCausalLM, so mainstream engines load it with no custom kernels and no model-code fork. Is it deployable? Yes. The weights are Apache 2.0 and run through vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX and FlagOS. What the benchmark table actually shows OpenBMB compares MiniCPM5-2B against LFM2.5-2.6B, Qwen3.5-2B and Gemma-4-E2B-it in the same size class, and lists Qwen3.5-4B, granite-4.2-3B, Nemotron-3-Nano-4B, Gemma-4-E4B-it and LFM2.5-8B-A1B for reference. Across 34 benchmark rows it averages 53.9. The best baseline in that set is Qwen3.5-4B at 51.1, then granite-4.2-3B at 42.7 and LFM2.5-2.6B at 33.2. On code reasoning MiniCPM5-2B posts 69.1 on LiveCodeBench v6 against 56.4, and 46.4 on SWE-bench Verified against 33.6. Tool use is the widest margin: 97.1 on τ²-Bench Telecom, 66.6 on BFCL v4, and 20.8 on τ³-Bench Banking against 6.8. Long context is split, with 68.1 on NoLiMa against 43.5, but 59.0 on AA-LCR against 61.0 and 43.7 on LongBench v2 against 47.3. General knowledge is where the size gap shows: 70.8 on MMLU-Pro against 78.0, and 8.9 on Humanity’s Last Exam against 9.9. OpenBMB marks rows sourced from Artificial Analysis separately from internally reproduced ones. Training recipe: SFT, then RL, then on-policy distillation Training follows the UltraData tiered data management method described in original research. Base training runs stable and decay phases, then mid-training adapts the model to the target data distribution. Post-training starts with 400B tokens of deep-thinking SFT, then trains specialised RL teachers for math, code, agentic tasks and writing using the critic-based JustRL II algorithm. The final step is on-policy distillation. OPD merges 16 RL experts, five of them agentic, into a single shipped model. At each response position it computes full-vocabulary reverse KL divergence between student and teacher logits as the advantage estimate, replacing the verification-based advantage. It reuses the RL prompts as distillation data, so no new corpus is built. OpenBMB measures the RL plus OPD stage at 10.96 average points on reasoning and general benchmarks and 6.96 points on agentic ones. The data is open too Alongside the weights, OpenBMB released Ultra-FineWeb, Ultra-FineWeb-L3, UltraX, UltraData-Code, UltraData-Math, UltraData-SFT-2605, UltraData-SFT-Agent-2609 with 500K agent samples, and UltraData-RL-2609 with more than 80K RL samples. Intermediate checkpoints are published as well, covering Base, Midtrain and SFT-only, so the contribution of each stage can be measured directly. Summary MiniCPM5-2B is a credible on-device option for agentic and tool-calling workloads, not a general knowledge model. Its advantage is clearest on tool use, coding agents and NoLiMa-style long-context retrieval, and it trails larger models on MMLU-Pro, GPQA-Diamond and MATH-500. The open data and intermediate checkpoints make the RL plus OPD claim checkable, which matters more than the headline average. Key Takeaways 2.52B dense model, 131,072 token context, Apache 2.0, standard Llama architecture. Averages 53.9 across 34 benchmarks, ahead of Qwen3.5-4B at 51.1. Strongest on tool use, coding agents and long-context retrieval; weakest on knowledge. Post-training pairs 400B SFT tokens with RL teachers and on-policy distillation. Pre-training, SFT and RL datasets ship alongside the weights. Check out the HF, GitHub repo and Web. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device appeared first on MarkTechPost.

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device Read Post »

AI, Committee, 新闻, Uncategorized

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

A team of researchers from UC Berkeley have released CUA-Lite, an open platform for computer-use agents (CUAs). The argument behind it is infrastructural rather than model-centric: training and benchmarking a CUA requires four pieces: agents, environments, traces, and a framework to evaluate and train them and all four are currently fragmented across separate repositories with incompatible interfaces. CUA-Lite puts them behind one action space, one data schema, and one command, across desktop, browser and mobile. Is it deployable? Yes. The stack installs with uv sync –all-extras on Python 3.12, and its lightweight sandboxes run on any Docker host without /dev/kvm, so cloud instances, CI runners and nested containers all work. The VM tax, and how Lite.OSWorld removes it The most concrete contribution is Lite.OSWorld. OSWorld provides a faithful Ubuntu desktop, but it ships as a full QEMU/KVM virtual machine per task, requiring nested virtualization that most managed infrastructure does not expose. CUA-Lite reproduces the same task suite and the same evaluators on a GNOME desktop inside a plain Docker container. Task OSWorld Lite.OSWorld Runtime QEMU/KVM VM Docker container Host requirement /dev/kvm, nested virt Any Docker host Memory 4.1 GB 0.9 GB Cold start 29.9 s 23.8 s Parallelism baseline ~4.6× more instances Task suite OSWorld Identical Fidelity is the obvious concern when you swap a VM for a container, and the team addresses it directly: across 13 models, Lite.OSWorld scores match the OSWorld VM’s, so a score or a training signal earned in the container transfers back to the real benchmark. The same base now carries a family of sandboxes: Lite.ScaleCUA, Lite.CUAGym and Lite.CUAWorld, the last expanding into roughly 40 applications including Blender, QGIS and VS Code. In total the platform claims 30k+ verifiable tasks. One schema for data, one adapter per model CUA-Lite’s second layer is LiteSample, a single supervised-learning schema shared across every environment, agent and task type, shipped as plain parquet plus images. Ten-plus existing CUA datasets have been preprocessed into it and published free on Hugging Face, including Aguvis, OpenCUA, ScaleCUA, GUI-360, GUIOdyssey and Multimodal-Mind2Web. Alongside those corpora sit fresh rollout datasets generated by rolling a frontier teacher model through the sandboxes, for distillation into smaller students. Because model families expect different scaffolding, the framework ships a per-model adapter that packs a unified LiteSample into each model’s own training format, including history collapsing so several steps share one forward pass. Eval, SFT and RL behind one command Agents and environments meet in lite.gym: screenshots up, actions down, with one action space per platform. 10+ agents are built in GPT, Claude, Gemini, Qwen3-VL, UI-TARS, Fara-7B, MAI-UI and others, and 15+ benchmarks are integrated, spanning grounding (ScreenSpot-Pro, OSWorld-G), desktop (OSWorld, OSWorld-2, WindowsAgentArena, CUABench), browser (WebArena, VisualWebArena, MiniWoB, WebVoyager, Online-Mind2Web, WebGym) and mobile (AndroidWorld, AndroidLab, MobileWorld, MobileGym). Swapping –model-id and –env-id in scripts/rollout.py is the whole interface. The same loop serves training. For SFT, the README documents fine-tuning Qwen3-VL-2B-Instruct on Lite.ScaleCUA desktop trajectories, lifting mean episode return from 0.138 to 0.237 on the 332-task lite.osworld eval split, a single reported configuration on two GPUs, not an independently reproduced result. For RL, rollouts scored in the environment drive GRPO updates on top of Slime, with a worked MobileGym example covering 416 mobile tasks across 28 apps. Interactive explainer Key Takeaways CUA-Lite unifies agents, environments, traces and training under one action space and one LiteSample schema. Lite.OSWorld runs OSWorld tasks VM-free in Docker at 0.9 GB versus 4.1 GB, roughly 4.6× more parallel desktops. Scores in the container match the OSWorld VM across 13 models, so training signal transfers to the real benchmark. 30k+ verifiable tasks, 15+ benchmarks, 10+ agents, and 20+ datasets published free on Hugging Face. Deployable on any Docker host, but the repository ships no explicit license yet — verify terms before commercial use. Check out the Project Page, GitHub Repo and Datasets on Hugging Face. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents appeared first on MarkTechPost.

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents Read Post »

AI, Committee, 新闻, Uncategorized

Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second — the serving infrastructure behind pplx-embed and the ranking models used across Perplexity Search, Computer and the API Platform. Perplexity team states that embedding inference on the GPU side has largely converged across engines on mature Hopper and Blackwell hardware. The wins sit in the runtime and harness around the model: CUDA graph management, an async result-tracking abstraction, and a Rust request path. Two traffic patterns, one engine Perplexity frames embedding serving as two workloads. Batch embedding happens when building or re-indexing the vector database, where throughput minimizes cost. Online embedding happens at query time, where a short query must be embedded fast. Scoring sits in between: after vector search, large document batches are ranked, balancing both. The key decision is that Perplexity did not build a separate embedding engine. Because embedding models are small Transformers, batch embedding resembles compute-bound prefill and online embedding, often a few tokens, resembles memory-bound decode. So the research team reuses the prefill and decode kernels from its LLM stack. Ivy, Tulip and ROSE Three services handle a request: Ivy is a Rust HTTP gateway. It does the CPU-side work — JSON parsing, tokenization, input templating, batch splitting — and translates requests into a custom gRPC protocol. It also splits large-batch requests into chunks and load-balances them across replicas, which corrects the load imbalance that arises when production payloads vary in size. Tulip is the inference server interface: a gRPC server built with Rust, tokio and tonic, handling scheduling and batching before dispatching to the engine. ROSE (Runtime-Optimized Serving Engine) implements model inference. It is primarily Python, provides kernels, layers and model definitions, manages CUDA graphs, and exposes a step() function to Tulip. Why the scheduler is deliberately simple Tulip picks sequences first-come, first-served while requests accumulate. That simplicity is justified by a measurement: for small embedding models at the sequence lengths Perplexity serves, the linear cost of dense layers dominates the quadratic cost of attention. Latency is therefore roughly proportional to token count, not sequence count. Once a batch saturates the GPU, around 512 tokens on a sub-billion-parameter model, packing in more sequences does not improve efficiency. CUDA graphs and LazyTensors On small batches, CPU-side kernel launching can outweigh GPU execution. Perplexity builds whole-model CUDA graphs for all embedding models, capturing every launch into a single driver call. Because embedding models are small, the inflection point where GPU work exceeds launch cost arrives at batches of thousands of tokens and tens of sequences. Some attention implementations block full-model graphs by depending on dynamic host-side inputs; Perplexity upstreamed changes to FlashInfer to enable capture. Graphs must be captured per configuration, so token counts are padded to buckets that are multiples of 64 or 256. That still yields thousands of graphs and multiple minutes of capture per model. The fix is lazy capture: each configuration gets an eager warmup run, then triggers capture and replay on its second hit. This costs p99 latency at startup but spreads minutes of eager work across hours. The second piece is the LazyTensor, which tracks a page-locked host buffer plus a cudaMemcpyAsync and a CUDA event. Instead of step() blocking on the device, it returns a LazyTensor, letting a Rust async task wait on batch N while the CPU enqueues N+1. Send a request</button></div> </div> </div> <div class="”pe-panel”" id="”peP1″"> <div class="”pe-card”"> <div class="”pe-txt”">On small batches, CPU-side kernel launches can outweigh GPU work. A whole-model <b>CUDA graph</b> captures every launch into one call to the driver, so the CPU is freed to enqueue the next batch. Toggle the two modes.</div> <div class="”pe-ctl”" style="”margin:0" 0 12px”> <button class="”pe-btn" ghost on” id="”peEager”">Eager launches</button> <button class="”pe-btn" ghost” id="”peGraph”">CUDA graph</button> </div> <div class="”pe-lane”"><div class="”pe-lbl”">Host / CPU</div><div class="”pe-track”" id="”peCpuT”"></div></div> <div class="”pe-lane”"><div class="”pe-lbl”">Device / GPU</div><div class="”pe-track”" id="”peGpuT”"></div></div> <div class="”pe-stats”"> <div class="”pe-stat”"><div class="”v”" id="”peLaunches”">—</div><div class="”k”">Driver calls</div></div> <div class="”pe-stat”"><div class="”v”" id="”peGap”">—</div><div class="”k”">GPU idle gaps</div></div> </div> <div class="”pe-txt”" style="”margin:12px" 0 0;font-size:11.5px;color:#6e8285″>Schematic. Block widths illustrate the launch-overhead pattern described in the post, not measured timings.</div> </div> </div> <div class="”pe-panel”" id="”peP2″"> <div class="”pe-card”"> <div class="”pe-txt”">Reading results back normally forces a host sync. A <b>LazyTensor</b> tracks a page-locked host buffer plus an async device-to-host copy and a CUDA event, so Tulip can block on batch N while the CPU already prepares batch N+1.</div> <div class="”pe-lane”"><div class="”pe-lbl”">CPU — prepare / sync</div><div class="”pe-track”" id="”peLzC”"></div></div> <div class="”pe-lane”"><div class="”pe-lbl”">GPU — forward pass</div><div class="”pe-track”" id="”peLzG”"></div></div> <div class="”pe-ctl”"> <button class="”pe-btn”" id="”peLzRun”"> Run 3 batches</button> <button class="”pe-btn" ghost on” id="”peLzOn”">Overlapped</button> <button class="”pe-btn" ghost” id="”peLzOff”">Blocking</button> </div> <div class="”pe-note”" style="”margin-top:12px”" id="”peLzNote”">Overlapped: while the GPU chews batch N, the CPU is already tokenizing and packing batch N+1.</div> </div> </div> <div class="”pe-panel”" id="”peP3″"> <div class="”pe-card”"> <div class="”pe-txt”">For small embedding models at these sequence lengths, the linear cost of dense layers dominates the quadratic cost of attention, so latency tracks <b>token count, not sequence count</b>. Past roughly <b>512 tokens</b> on a sub-1B model, the GPU is saturated and packing in more sequences stops helping.</div> <div class="”pe-lbl”" style="”margin-top:6px”">Tokens in batch: <span id="”peTokV”" style="”color:#3FB6C4″">512</span></div> <input type="”range”" id="”peTok”" min="”32″" max="”4096″" step="”32″" value="”512″"> <div class="”pe-lbl”">GPU utilisation</div> <div class="”pe-bar”"><div class="”pe-fill”" id="”peUtil”"></div></div> <div class="”pe-stats”"> <div class="”pe-stat”"><div class="”v”" id="”peUtilV”">—</div><div class="”k”">Saturation</div></div> <div class="”pe-stat”"><div class="”v”" id="”peState”">—</div><div class="”k”">Regime</div></div> </div> <div class="”pe-txt”" style="”margin:12px" 0 0;font-size:11.5px;color:#6e8285″>Illustrative curve. The ~512-token saturation point is the figure stated in the post; the shape between points is a stand-in, not a benchmark.</div> </div> </div> <div class="”pe-foot”"> <span>Source: Perplexity Engineering, “Fast Embeddings on GPUs” (Sep 4, 2026)</span> <span><a href="/zh/”https://www.marktechpost.com”/">Built by Marktechpost</a></span> </div> <script> (function(){ var R=document.getElementById(‘pplxEmbedExplainer’); var NOTES=[ ‘<b>Ivy</b> — parses JSON, tokenizes with the in-house unigram tokenizer, applies input templating and splits large batches, then translates to a custom gRPC protocol. It also load-balances chunks across replicas.’, ‘<b>Tulip</b> — Rust gRPC server on tokio and tonic. Requests accumulate while it dispatches or waits; sequences are picked first-come, first-served and packed into a batch for the accelerator.’,

Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed Read Post »

AI, Committee, 新闻, Uncategorized

Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours

AI research agents can already propose, implement and score their own machine learning experiments. Idea generation is cheap; verification is not. Training one candidate can consume hours to days of GPU time, so an agent proposes far more candidates than it can afford to run. Which ones get run is the real lever on research progress. A research team from FAIR at Meta, the University of Oxford and University College London formalizes that lever as research preference and introduces AI Research Preference Models (RPMs). An RPM ranks unexecuted candidates and picks one to execute. It never forecasts an absolute score, the team found language models unreliable at predicting metrics or execution outcomes. Is it deployable? Partially. RPMs use frozen pretrained LLMs with no fine-tuning, the scaffold AIRA-dojo and benchmark AIRS-Bench are open source, and the backbone Qwen3.6-27B is open weights. Where the RPM sits in the agent loop AIRA-dojo is an evolutionary tree search: greedy parent selection, Draft / Improve / Debug operators, highest-validation-score node returned at the end. The RPM intervenes at child creation only. Instead of generating one child and executing it, the agent applies the operator 15 times in parallel to yield 15 unexecuted candidates, then compares them pairwise in a knockout tournament. Only the winner is executed. Each comparison is grounded in context nodes collected by a BFS walk of the explored tree, each shown with the validation score it obtained. Two variants, two compute budgets Inference-only RPM: An LLM-as-a-judge over candidate plans, code and search history. Its prompt was optimized with MIPROv2 from DSPy, converging on a principal-investigator rubric that tolerates fixable bugs, rewards extensibility and penalizes redundant directions, offline accuracy 57.7% to 59.0%. Agentic RPM: The same judge and a sandbox that clones the agent’s environment, including a single H200. Tools are python, bash and submit_solution. It runs small-scale pilot experiments, then a feedback model either proposes the most informative next experiment or ends the loop. Two design choices carry weight: the remaining budget is deliberately overstated (2,700s reported against a real 300s) so the agent does not stop early, and pilots are capped at 30 with a 60-second threshold. Pilot time competes with the agent’s own clock, so the agentic selector runs only on Draft and Improve steps; Debug reverts to random. Results on AIRS-Bench Setup: 20 public text and tabular tasks, 24 hours on a single H200 per task, 10 seeds, Qwen3.6-27B as backbone for both the operators and the RPM, so the gain comes from the selection layer, not a stronger judge. Child selection Avg. normalized score No RPM (random pick) 0.684 Inference-only RPM 0.711 Agentic RPM 0.729 Validation oracle (ceiling) 0.748 Test oracle (ceiling) 0.759 Probability of improvement over No-RPM is 0.5923 and 0.5913, with 95% CI lower bounds at 0.5066 and 0.5018. Efficiency is the more practical result. Inference-only reaches the baseline’s final 0.684 in 14.88 hours (1.61×), agentic in 15.50 hours (1.55×). Self-hosted inference adds 0.660 hours per run; adjusting for it still gives 0.708 at 23.34 hours. Two new reported SOTA results: WinoGrande 94.1% with the Agentic RPM against a prior agentic SOTA of 90.4% from AIRA₂, and SVAMP 95.7% with inference-only against a prior human SOTA of 94.2%. Key Takeaways RPMs rank unexecuted candidates so an AI research agent runs only the most promising one. Two frozen-LLM variants: an inference-only judge, and an agentic judge that runs short pilots. On AIRS-Bench, average normalized score rises from 0.684 to 0.711 and 0.729. Both hit the baseline’s 24-hour score in roughly 15 hours, a 1.5–1.6× speedup. New reported SOTA on WinoGrande (94.1%) and SVAMP (95.7%). Check out the Paper and the LinkedIn announcement. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours appeared first on MarkTechPost.

Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours Read Post »

AI, Committee, 新闻, Uncategorized

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

Most visual document retrievers in production today are hand-me-downs. ColPali and the models that followed it take a generative vision-language model and repurpose it as an encoder. The result still carries a separately pretrained vision tower and a causal decoder that never generates a token. That is parameter and compute overhead for a task that only needs representations. H Company has released NeoMME, a family of 260M and 800M bidirectional encoders that drops both components. One Transformer processes multilingual text tokens and raw 32×32 RGB image patches through the same layers, trained from random initialization. The retrieval fine-tune, NeoMME-Retriever, reaches 0.523 nDCG@10 on ViDoRe v3 at 260M parameters. Is it deployable? Yes. Every checkpoint ships under Apache 2.0 with day-zero support in Hugging Face Transformers. The 260M model indexes 51.3 pages per second on a single NVIDIA L40S and encodes a query in 78.3 ms on a CPU-only host. https://arxiv.org/pdf/2609.01657 One tower, two modalities Text enters through an ALBERT-style factorized embedding: a 256-dimensional lookup projected to model width. Images are split into non-overlapping 32×32 patches and projected by a 2-layer MLP trained from scratch. No patch-merging module, no SigLIP2 tower. Both models support a 16,384-token context, enough for two standard 3,840×2,160 4K UHD images after patching. Most layers use symmetric sliding-window attention; every sixth layer and the final layer attend globally. The stack uses grouped-query attention, query-key normalization, gated attention, 2D rotary position embeddings, and squared-ReLU MLPs. Exact parameter counts are 262,937,906 and 793,715,032. The tokenizer is a whitespace-unconstrained BPE with a 131,072-entry vocabulary, trained from scratch. Across 14 target languages in FLORES-200 devtest, it emits 44.4% fewer tokens than ModernBERT. Trained as a masked diffusion denoiser Pretraining is discrete masked diffusion over text, optionally conditioned on visible image patches. Text-only segments draw a corruption rate uniformly from 0 to 1. Multimodal segments draw from 0.30 to 1, which removes the language-only shortcut and forces the model to read the page. A cross-modal ablation probe confirms this works. At 90% masking, visible page patches raise masked-token accuracy by 38.4 points for the 260M model and 40.5 points for the 800M model. Each run processes about 524 billion packed input tokens, roughly 290 billion of them text-only, on 16 and 32 H100 accelerators respectively. Retrieval results NeoMME-Retriever adds two jointly trained heads on the shared backbone: a mean-pooled dense head with Matryoshka widths, and a late-interaction head projecting every token and patch to 128 dimensions. One forward pass returns both. On ViDoRe v3, the 260M model scores 0.523 nDCG@10 and the 800M model 0.556. The 260M result sits within 0.002 of ColQwen2.5-v0.2 at 3.75B parameters, and 26.1 points above the best other sub-300M model. The 800M model lands 0.9 points behind the similarly sized Vultron Retriever Flash. On ViDoRe v1 and v2 the models reach 0.860/0.522 and 0.874/0.559 nDCG@5. Text retrieval is weaker. On BEIR-15, late interaction reaches 0.4881 and 0.5126, against 0.5722 for LateOn at 149M parameters. The authors attribute this partly to supervision scale: NeoMME saw roughly 430K text query examples, against roughly 660M contrastive examples for mLateOn. Storage and throughput Late-interaction indexes are expensive. A 2048×2048 page yields 4,162 vectors, about 1.5 MB per ViDoRe v3 document in float32. Two methods bring that down. Hierarchical token pooling at factor 10 with int8 queries and documents gives 39.0 kB per page, a 39.4× reduction retaining 99.16% of baseline nDCG@10. Pool factor 8 with int8 queries and binary documents gives 6.0 kB, a 255.5× reduction retaining 95.19%. Indexing is fast for the vector count. At a matched 2048×2048 input on one L40S, NeoMME-260M encodes 51.3 pages per second against ColModernVBERT’s 26.0, a 1.97× gap. Interactive explainer Key Takeaways One bidirectional Transformer handles text and raw image patches, with no vision tower and no decoder. NeoMME-Retriever-260M scores 0.523 nDCG@10 on ViDoRe v3, beating every evaluated model below 800M. It matches 3.75B-parameter ColQwen2.5 on ViDoRe v3 while being 14.4× smaller. Token pooling plus asymmetric quantization cut the index from roughly 1.5 MB to 6 kB per page. Text-only retrieval and frozen natural-image transfer remain clear weak spots. Check out the Paper, Model Collection and Demo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder appeared first on MarkTechPost.

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder Read Post »

We use cookies to improve your experience and performance on our website. You can learn more at 隱私權政策 and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
zh_CN