YouZum

Uncategorized

AI, Committee, Notizie, Uncategorized

Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration

arXiv:2602.21368v1 Announce Type: cross Abstract: Given a black-box AI system and a task, at what confidence level can a practitioner trust the system’s output? We answer with a reliability level — a single number per system-task pair, derived from self-consistency sampling and conformal calibration, that serves as a black-box deployment gate with exact, finite-sample, distribution-free guarantees. Self-consistency sampling reduces uncertainty exponentially; conformal calibration guarantees correctness within 1/(n+1) of the target level, regardless of the system’s errors — made transparently visible through larger answer sets for harder questions. Weaker models earn lower reliability levels (not accuracy — see Definition 2.4): GPT-4.1 earns 94.6% on GSM8K and 96.8% on TruthfulQA, while GPT-4.1-nano earns 89.8% on GSM8K and 66.5% on MMLU. We validate across five benchmarks, five models from three families, and both synthetic and real data. Conditional coverage on solvable items exceeds 0.93 across all configurations; sequential stopping reduces API costs by around 50%.

Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Nous Research Releases ‘Hermes Agent’ to Fix AI Forgetfulness with Multi-Level Memory and Dedicated Remote Terminal Access Support

In the current AI landscape, we’ve become accustomed to the ‘ephemeral agent’—a brilliant but forgetful assistant that restarts its cognitive clock with every new chat session. While LLMs have become master coders, they lack the persistent state required to function as true teammates. Nous Research team released Hermes Agent, an open-source autonomous system designed to solve the two biggest bottlenecks in agentic workflows: memory decay and environmental isolation. Built on the high-steerability Hermes-3 model family, Hermes Agent is billed as the assistant that ‘grows with you.’ The Memory Hierarchy: Learning via Skill Documents For an agent to ‘grow,’ it needs more than just a large context window. Hermes Agent utilizes a multi-level memory system that mimics procedural learning. While it handles short-term tasks through standard inference, its long-term utility is driven by Skill Documents. When Hermes Agent completes a complex task—such as debugging a specific microservice or optimizing a data pipeline—it can synthesize that experience into a permanent record. These records are stored as searchable markdown files following the agentskills.io open standard. Procedural Memory: The next time you ask the agent to perform a similar task, it doesn’t start from scratch. It queries its own library of Skill Documents to ‘remember’ the successful steps it took previously. Contextual Persistence: Unlike standard RAG (Retrieval-Augmented Generation), which often pulls disjointed snippets, this system allows the agent to maintain a cohesive understanding of your specific codebase and preferences over weeks or months. Persistent Machine Access: Beyond the Sandbox A major friction point for AI devs is the ‘execution gap.’ Most agents write code but cannot interact with the real world without heavy manual intervention. Hermes Agent closes this gap by providing persistent dedicated machine access. The agent is designed to live inside a functional environment, supporting five distinct backends: Local: Direct interaction with the host machine. Docker: Isolated, reproducible containers for safe code execution. SSH: The ability to log into remote servers or cloud instances. Singularity: High-performance computing (HPC) container support. Modal: Serverless execution for scaling heavy workloads. This persistence is critical for AI devs. You can initialize a long-running EDA (Exploratory Data Analysis) on a remote server via SSH, log-off, and return later. The agent maintains the terminal state, handles background processes, and tracks file system changes independently. It isn’t just simulating a conversation; it is managing a workspace. The Gateway: An Agent in Your Pocket While most technical agents are confined to a CLI or a proprietary web dashboard, Nous Research has prioritized accessibility through the Hermes Gateway. The system integrates directly with existing communication stacks, including Telegram, Discord, Slack, and WhatsApp. This allows for a continuous feedback loop: an engineer can start a task at their workstation and receive a ‘task completed’ notification via Telegram. Through the gateway, you can send follow-up instructions or even voice memos that the agent processes and executes within its persistent environment. Under the Hood: The ReAct Loop and Steerability For the AI devs building on this, the architecture is a refined implementation of the ReAct (Reasoning and Acting) loop. The agent follows a structured cycle: Observation: Reading terminal output or file contents. Reasoning: Analyzing the current state against the goal. Action: Executing a command or calling a tool. This is powered by Hermes-3 (based on Llama 3.1), which was trained using a specialized reinforcement learning framework called Atropos. This training specifically targets tool-calling accuracy and long-range planning, ensuring the agent doesn’t get ‘lost’ during multi-step deployments. Key Takeaways Persistent Machine Access: Unlike stateless chatbots, it operates in real terminal environments (Docker, SSH, Local, etc.), allowing it to run long-term tasks and maintain file states across sessions. Self-Evolving ‘Skill Documents’: It uses a multi-level memory system to record successful workflows as searchable markdown files (via agentskills.io), meaning it literally gets smarter the more you use it. Precision ‘Hermes-3’ Thinking: Powered by the Llama 3.1-based Hermes-3 model, it is fine-tuned with Atropos RL for high steerability and reliable tool-calling within complex reasoning loops. Omnipresent Gateway: You can interact with your agent via Telegram, Discord, or Slack, enabling you to manage heavy engineering tasks or receive status updates from your phone. Check out the Technical details and GitHub Repo. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Nous Research Releases ‘Hermes Agent’ to Fix AI Forgetfulness with Multi-Level Memory and Dedicated Remote Terminal Access Support appeared first on MarkTechPost.

Nous Research Releases ‘Hermes Agent’ to Fix AI Forgetfulness with Multi-Level Memory and Dedicated Remote Terminal Access Support Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

America was winning the race to find Martian life. Then China jumped in.

To most people, rocks are just rocks. To geologists, they are much, much more: crystal-filled time capsules with the power to reveal the state of the planet at the very moment they were forged.  For decades, NASA had been on a time capsule hunt like none other—one across Mars. Its rovers have journeyed around a nightmarish ocher desert that, billions of years ago, was home to rivers, lakes, perhaps even seas and oceans. They’ve been seeking to answer a momentous question: Once upon a time, did microbial life wriggle across its surface?  Then, in July 2024, after more than three years on the planet, the Perseverance rover came across a peculiar rocky outcrop. Instead of the usual crystals or layers of sediment, this one had spots. Two kinds, in fact: one that looked like poppy seeds, and another that resembled those on a leopard. It’s possible that run-of-the-mill chemical reactions could have cooked up these odd features. But on Earth, these marks are almost always produced by microbial life. To put it plainly: Holy crap. Sure, those specks are not definitive proof of alien life. But they are the best hint yet that life may not be a one-off event in the cosmos. And they meant the most existential question of all—Are we alone?—might soon be addressed. “If you do it, then human history is never the same,” says Casey Dreier, chief of space policy at the Planetary Society, a nonprofit that promotes planetary exploration and defense and the search for extraterrestrial life. But the only way to confirm whether these seeds and spots are the fossilized imprint of alien biology is to bring a sample of that rock home to study.  Perseverance was the first stage of an ambitious scheme to do just that—in effect, to pull off a space heist. The mission—called Mars Sample Return and planned by the US, along with its European partners—would send a Rube Goldberg–like series of robotic missions to the planet to capture pristine rocks. The rover’s job was to find the most promising stones and extract samples; then it would pass them to another robot—the getaway driver—to take them off Mars and deliver them to Earth. But now, just over a year and a half later, the project is on life support, with zero funding flowing in 2026 and little backing left in Congress. As a result, those oh-so-promising rocks may be stuck out there forever. “We’ve spent 50 years preparing to get these samples back. We’re ready to do that,” says Philip Christensen, a planetary scientist at Arizona State University who works closely with NASA. “Now we’re two feet from the finish line—Oh, sorry, we’re not going to complete the job.” This also means that, in the race to find evidence of alien life, America has effectively ceded its pole position to its greatest geopolitical rival: China. The superpower is moving full steam ahead with its own version of MSR. It’s leaner than America and Europe’s mission, and the rock samples it will snatch from Mars will likely not be as high quality. But that won’t be the headline people remember—the one in the scientific journals and the history books. “At the rate we’re going, there’s a very good chance they’ll do it before we do,” laments Christensen. “Being there first is what matters.”   Of course, any finding of extraterrestrial life advances human knowledge writ large, no matter the identity of the discoverers. But there is the not-so-small issue of pride in an already heated nationalistic competition, not to mention the fact that many scientists in America (to say nothing of US lawmakers) don’t necessarily want their future research and scientific progress subject to a foreign gatekeeper. And even for those not especially concerned about potentially unearthing alien microbes, MSR and the comparable Chinese mission are technological stepping stones toward a long-held dream shared by many beyond Elon Musk: getting astronauts onto the Red Planet and, eventually, setting up long-term bases for astronauts there. It’d be a huge blow to show up only after a competitor had already set up shop … or not to get there at all.  “If we can’t do this, how do we think we’re gonna send humans there and get back safely?” says Victoria Hamilton, a planetary geologist at the Southwest Research Institute in Boulder, Colorado, who is also the chair of the NASA-affiliated Mars Exploration Program Analysis Group.  Or as Paul Byrne, a planetary scientist from the Washington University in St. Louis, puts it: “If you’re going to bring humans back from Mars, you sure as shit have to figure out how to bring the samples back first.”  Nearly a dozen project insiders and scientists in both the US and China shared with me the story of how America blew its lead in the new space race. It’s full of wild dreams and promising discoveries—as well as mismanagement, eye-watering costs, and, ultimately, anger and disappointment.     “I spent most of my career studying Mars,” says Christensen. There are countless things about it that bewitch him. But by examining it, he suspects, we’ll get further than ever in the Homeric investigation of how life began. Sure, the Mars of today is a postapocalyptic wasteland, an arid and cold desert bathed in lethal radiation. But billions of years ago, water lapped up against the slopes of fiery volcanoes that erupted under a clement sky. Then its geologic interior cooled down so quickly, changing everything. Its global magnetic field collapsed like a deflating balloon, and its protective atmosphere was stripped away by the sun.  NASA first touched down on Mars in 1976 with two Viking landers. The Mars Odyssey spacecraft has been orbiting the planet since 2001 and produced this image of Valles Marineris, which is 10 times longer, 5 times deeper, and 20 times wider than the Grand Canyon. NASA/ARIZONA STATE UNIVERSITY VIA GETTY IMAGES Its surface is now remarkably hostile to life as we know it. But deep below ground, where it’s shielded from space, and

America was winning the race to find Martian life. Then China jumped in. Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

This company claims a battery breakthrough. Now they need to prove it.

When a company claims to have created what’s essentially the holy grail of batteries, there are bound to be some questions. Interest has been swirling since Donut Lab, a Finnish company, announced last month that it had a new solid-state battery technology, one that was ready for large-scale production. The company said its batteries can charge super-fast and have a high energy density that would translate to ultra-long-range EVs. What’s more, it claimed the cells can operate safely in the extreme heat and cold, contain “green and abundant materials,” and would cost less than lithium-ion batteries do today. It sounded amazing—this sort of technology could transform the EV industry. But many quickly wondered if it was all too good to be true. Now, Donut Lab is releasing a series of videos that it says will prove its technology has the secret sauce. Let’s dig into why this company is making news, why many experts are skeptical, and what it all means for the battery industry right now. Solid-state batteries could deliver the next generation of EVs. In place of a liquid electrolyte (the material that ions move through inside a battery), the cells use a solid material, so they can be more compact. That means a significantly longer range, which could get more people excited to drive EVs. The problem is, getting these batteries to work and making them at the large scale required for the EV industry hasn’t been a simple task. Some of the world’s most powerful automakers and battery companies have been trying for years to get the technology off the ground. (Toyota at one point said it would have solid-state batteries in cars by 2020. Now it’s shooting for 2027 or 2028.) While it’s been a long time coming, it does feel as if solid-state batteries are closer than ever. Much of the progress so far has been on semi-solid-state batteries, which use materials like gels for electrolytes. But some companies, including several in China, are getting closer to true solid state. The world’s largest battery company, CATL, plans to manufacture small quantities in 2027. Another major Chinese automaker, Changan, plans to start testing installation of all-solid-state batteries in vehicles this year, with mass production expected to begin next year. Still, Donut Lab surprised the battery industry when, in a video released in early January ahead of the Consumer Electronics Show in Las Vegas, the company claimed it would put the world’s first all-solid-state battery into production vehicles. One of the splashiest claims in the announcement was that cells would have an energy density of 400 watt-hours per kilogram (the top commercial lithium-ion batteries today sit at about 250 to 300 Wh/kg). It was also claimed that the cells could charge in as little as five minutes, last 100,000 cycles, and retain 99% of capacity at high and low temperatures—while costing less than lithium-ion cells and being made from “100% green and abundant materials with global availability.” Many experts were immediately skeptical. “In the solid-state field, the technical barriers are very high,” said Shirley Meng, a professor of molecular engineering at the University of Chicago, when I spoke with her last month. She’d recently attended CES and visited Donut Lab’s booth. “They had zero demo, so I don’t believe it,” she says. “Call me conservative, but I would rather be careful than be sorry later.” “It’s one of those things where nobody knows—they’ve never heard of it,” said Eric Wachsman, a professor at the University of Maryland and cofounder of the solid-state battery company Ion Storage Systems, in a January interview. “They came out of nowhere.” Donut Lab has shared very little about what, exactly, this technology might be. It’s not uncommon for battery companies (or any startup, for that matter) to be quiet about technical details before they can get patents filed to protect their technology. But the combination of claims didn’t seem to line up with any known chemistries, leaving experts speculating and, in many cases, doubting Donut Lab’s claims. “All the parameters are contradictory,” said Yang Hongxin, chairman and CEO of the Chinese battery giant Svolt Energy, in remarks to news outlets in January. For example, there’s often a trade-off between high energy density, which requires thicker electrodes that can store more energy, and fast charging, which requires ions to move quickly through cells. High-performance batteries are also expected to be costly, but Donut Lab claims its technology will be cheaper than lithium-ion technology.  In a new video released last week, Donut Lab cofounder and CEO Marko Lehtimäki announced the company would be releasing a video series, called “I Donut Believe,” that would provide evidence for their claims. As a header on the accompanying website reads: “Fair enough. Here you go.” When the website went up last week, it included a countdown timer to Monday February 23, when the company released results from its first third-party testing: a fast charging test. The test showed that a single cell could charge from 0% to 80% capacity in about four and a half minutes—incredibly quick and quite impressive results. (One potential caveat to note is that the cells heated up quite a bit, so thermal management could be important in designing vehicles that use these batteries.) Even as we see the first technical test results, I’m still left with a lot of questions. How many cycles could this battery do at this charging speed? Can this same cell meet the company’s other performance claims? (I’ve reached out to Donut Lab several times over the past month, both to the company’s press email and to leadership on LinkedIn, but I haven’t gotten a response yet.) The company has certainly drummed up a lot of interest and attention with its rollout, and the theatrics aren’t over yet. There’s another countdown timer on Donut Lab’s site, which ends on Monday, March 2. I’m the first one to get excited about a new battery technology. But there’s a sentiment I’ve seen pop up a lot recently online, and

This company claims a battery breakthrough. Now they need to prove it. Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Now is a good time for doing crime

Eons ago, in 2012, I had a weird experience. My iPhone suddenly shut down. When I restarted it, I found it was totally reset—clean, like a new device. This was the early days of iOS, so I wasn’t too concerned until I went to connect it to my computer to restore it from a backup. But when I flipped open the lid of my laptop, it too was mid-restart. And then, suddenly, the screen went gray. It was being remotely wiped. I turned on my iPad. It, too, had been wiped. I was being hacked.  Frantically, I shut down all my devices, unplugged everything connected to the internet in my house, turned off my router, and went next door to use my neighbors’ computer and find out what was going on. Deepening my panic, I realized hackers had also gained control of, and nuked, my Google account. Worse, they were in control of my Twitter, which they were gleefully using to spew all sorts of vile comments. It was nasty.  You have to remember, this was before all of us lived with a constant rain of text messages and emails designed to elicit the information necessary to pull something like this off. These crooks hadn’t brute-forced their way in, or used any sort of sophisticated techniques to gain access to my accounts. Instead, they had relied on publicly available information, and a fake credit card number, to socially engineer their way into my Amazon account, where they looked up the last four digits of my real credit card number. Then they used that information to get into Apple. And because that account was linked to my Gmail, and that to my Twitter, it gave them the keys to everything. But what really troubled me was what I learned as I followed up on my hack over the ensuing weeks and months: This kind of thing was, while still novel, becoming more common. Some version of what happened to me had happened to lots of other people. The kids who were responsible—it was a couple of kids—weren’t criminal masterminds. They had just found a gap, a place where a technology was now commonplace but its risks and exploitable surface areas weren’t yet fully understood. I just happened to have all my stuff in the gap. Today that gap might feature a crypto wallet or a deepfake of a loved one’s voice. (Or both.) Crime changes. The goals stay the same—pursuit of value, pursuit of power—but new technologies create new vulnerabilities, new tactics, and new ways for perpetrators to evade discovery or capture. And the law necessarily lags behind. Relying not on innovation but on precedent, it is intentionally backward-looking and slow. That plodding consideration used to be how we protected our shared democratic society, how we protected each other from each other. But those same new technologies that have allowed crime to outpace law have also reenergized law enforcement and government—offering new ways to root out crime, to gather evidence, to surveil people. Think, for example, of how cold-case investigators tracked down the Golden State Killer years after his murders, using DNA samples and genealogy databases—launching a new era of DNA-powered investigations.  Technology has long made crime and its prosecution a game of cat and mouse. It sometimes calls into question the nature of crime itself. Unregulated behaviors, facilitated by technology, can exist in murky zones of dubious legality. (Until TikTok announced its new ownership structure, Apple and Google were both technically breaking the law by allowing the app to stay on their platforms, under the provisions of the Protecting Americans from Foreign Adversary Controlled Applications Act. Ah! Well. Nevertheless.) That tension is the key to our March/April issue. Thanks to technologies like cryptocurrency and off-the-shelf autonomous autopilots, there’s never been a better time to do crime. Thanks to pervasive surveillance and digital infrastructure, there’s never been a better time to fight it—sometimes at the expense of what we used to think of as fundamental civil rights.  I never pressed charges against the kids who hacked me. The biggest consequence of the hack was that Apple set up two-factor authentication in the following months, which felt like a win. Now I’m not sure anyone expects their personal data to be secure in any meaningful way. I’m certain, though, that somewhere on the net, a new generation of kids is coming up with another novel crime. 

Now is a good time for doing crime Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

3 things Juliet Beauchamp is into right now

The only reality show that matters The Real Housewives of Salt Lake City is one of the best shows on television right now. Not one of the best reality TV shows, but one of the best TV shows, period. Chronicling a shifting group of wealthy women in and around Salt Lake, the show has featured a convicted felon whom federal agents came looking for while cameras were rolling, a church leader married to her step-grandfather, and a single mom in an exhausting on-again, off-again relationship with an Osmond. In one season, there was an ongoing argument between two cast members after one told the other that she “smelled like hospital.” Later, one woman was secretly running an anonymous gossip Instagram about her fellow housewives. We can debate the “reality” of reality television, and it’s certainly true that these characters and scenarios are far-fetched. But every single person is dealing with something relatable—difficult marriages, failing businesses, strained relationships with children, addiction. It’s entertainment, and high camp, but I find that I still have a lot of empathy for these people. The last good place(s) on Facebook Facebook sucks. That’s not controversial to say, right? But there is one reason I still have a Facebook account: my neighborhood Buy Nothing group. The spirit of community and camaraderie is alive and well there—and probably in yours, too. A non-exhaustive list of things I have given away: empty candle jars, a bookcase, used lightbulbs, unopened toiletries, bubble wrap. I’ve scored a few good things as well: a gorgeous antique dresser that I refinished, some over-the-door hooks, and brand-new jeans. It makes me happy to know that stuff that would’ve otherwise ended up in a landfill is bringing one of my neighbors joy. Going analog I used to wear an Apple Watch a lot. I’m a pretty active person, and I liked tracking my workouts and my steps. But after I’d had it for a while, my watch started dying in the middle of a 30-minute run; it became useless to me, and I gave it up completely. Guess what? I’m happier. I feel more present when I’m not checking how much time is left in a yoga class or reading texts during a long run. The amount of data it gathered about me was also stressing me out, and it wasn’t useful. And I don’t need a wearable to tell me how poorly I slept! Trust me, I already know.

3 things Juliet Beauchamp is into right now Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Listen to Earth’s rumbling, secret soundtrack

The boom of a calving glacier. The crackling rumble of a wildfire. The roar of a surging storm front. They’re the noises of the living Earth, music of this one particular sphere and clues to the true nature of these dramatic events. But as loud as all these things are, they emit even more acoustic energy below the threshold of human hearing, at frequencies of 20 hertz or lower. These “infrasounds” have such long wavelengths that they can travel around the globe as churning emanations of distant events. But humans have never been able to hear them. Until now, that is. Everyday Infrasound in an Uncertain World, a new album by the musician and artist Brian House, condenses 24 hours of these rumbles into 24 minutes of the most basic of bass lines, putting a new spin on the idea of ambient music. Sound, even infrasound, is really just variations in air pressure. So House built a set of three “macrophones,” tubes that funnel air into a barometer capable of taking readings 100 times a second. From the quiet woods of western Massachusetts, House can pick up what the planet is laying down. Then he speeds the recording up by a factor of 60 so that it’s audible to the wee ears of humans. “I am really interested in the layers of perception that we can’t access,” he says. “It’s not only low sound, but it’s also distant sound. That kind of blew my mind.” House’s album is art, but scientists made it possible. Barometers picked up the 1883 eruption of the South Pacific volcano Krakatoa as far away as London. And today, a global network of infrasound sensors helps enforce the nuclear test ban treaty. A few infrasound experts—like Leif Karlstrom, a volcanologist at the University of Oregon who uses infrasound to study Mount Kilauea in Hawaii—helped House set up his music-gathering array and better understand what he was hearing. “He’s highlighting interesting phenomena,” Karlstrom says, even though it’s impossible to tell exactly what is making each specific sound.  So how’s the actual music? It’s 24 minutes of an otherworldly chorus, alternating between low grumbling vibrations and soft ghostlike whispers. A high-pitched whistle? Could be a train, House says. An intense low-octave rattle? Maybe a distant thunderstorm or a shifting ocean current. “For me, it’s about the mystery of it,” he says. “I hope that’s a little bit unsettling.” But it also might connect someone listening to a wider—and deeper—world.  Monique Brouillette is a freelance writer based in Cambridge, Massachusetts.

Listen to Earth’s rumbling, secret soundtrack Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Liquid AI’s New LFM2-24B-A2B Hybrid Architecture Blends Attention with Convolutions to Solve the Scaling Bottlenecks of Modern LLMs

The generative AI race has long been a game of ‘bigger is better.’ But as the industry hits the limits of power consumption and memory bottlenecks, the conversation is shifting from raw parameter counts to architectural efficiency. Liquid AI team is leading this charge with the release of LFM2-24B-A2B, a 24-billion parameter model that redefines what we should expect from edge-capable AI. https://www.liquid.ai/blog/lfm2-24b-a2b The ‘A2B’ Architecture: A 1:3 Ratio for Efficiency The ‘A2B’ in the model’s name stands for Attention-to-Base. In a traditional Transformer, every layer uses Softmax Attention, which scales quadratically (O(N2)) with sequence length. This leads to massive KV (Key-Value) caches that devour VRAM. Liquid AI team bypasses this by using a hybrid structure. The ‘Base‘ layers are efficient gated short convolution blocks, while the ‘Attention‘ layers utilize Grouped Query Attention (GQA). In the LFM2-24B-A2B configuration, the model uses a 1:3 ratio: Total Layers: 40 Convolution Blocks: 30 Attention Blocks: 10 By interspersing a small number of GQA blocks with a majority of gated convolution layers, the model retains the high-resolution retrieval and reasoning of a Transformer while maintaining the fast prefill and low memory footprint of a linear-complexity model. Sparse MoE: 24B Intelligence on a 2B Budget The most important thing of LFM2-24B-A2B is its Mixture of Experts (MoE) design. While the model contains 24 billion parameters, it only activates 2.3 billion parameters per token. This is a game-changer for deployment. Because the active parameter path is so lean, the model can fit into 32GB of RAM. This means it can run locally on high-end consumer laptops, desktops with integrated GPUs (iGPUs), and dedicated NPUs without needing a data-center-grade A100. It effectively provides the knowledge density of a 24B model with the inference speed and energy efficiency of a 2B model. https://www.liquid.ai/blog/lfm2-24b-a2b Benchmarks: Punching Up Liquid AI team reports that the LFM2 family follows a predictable, log-linear scaling behavior. Despite its smaller active parameter count, the 24B-A2B model consistently outperforms larger rivals. Logic and Reasoning: In tests like GSM8K and MATH-500, it rivals dense models twice its size. Throughput: When benchmarked on a single NVIDIA H100 using vLLM, it reached 26.8K total tokens per second at 1,024 concurrent requests, significantly outpacing Snowflake’s gpt-oss-20b and Qwen3-30B-A3B. Long Context: The model features a 32k token context window, optimized for privacy-sensitive RAG (Retrieval-Augmented Generation) pipelines and local document analysis. Technical Cheat Sheet Property Specification Total Parameters 24 Billion Active Parameters 2.3 Billion Architecture Hybrid (Gated Conv + GQA) Layers 40 (30 Base / 10 Attention) Context Length 32,768 Tokens Training Data 17 Trillion Tokens License LFM Open License v1.0 Native Support llama.cpp, vLLM, SGLang, MLX Key Takeaways Hybrid ‘A2B’ Architecture: The model uses a 1:3 ratio of Grouped Query Attention (GQA) to Gated Short Convolutions. By utilizing linear-complexity ‘Base’ layers for 30 out of 40 layers, the model achieves much faster prefill and decode speeds with a significantly reduced memory footprint compared to traditional all-attention Transformers. Sparse MoE Efficiency: Despite having 24 billion total parameters, the model only activates 2.3 billion parameters per token. This ‘Sparse Mixture of Experts’ design allows it to deliver the reasoning depth of a large model while maintaining the inference latency and energy efficiency of a 2B-parameter model. True Edge Capability: Optimized via hardware-in-the-loop architecture search, the model is designed to fit in 32GB of RAM. This makes it fully deployable on consumer-grade hardware, including laptops with integrated GPUs and NPUs, without requiring expensive data-center infrastructure. State-of-the-Art Performance: LFM2-24B-A2B outperforms larger competitors like Qwen3-30B-A3B and Snowflake gpt-oss-20b in throughput. Benchmarks show it hits approximately 26.8K tokens per second on a single H100, showing near-linear scaling and high efficiency in long-context tasks up to its 32k token window. Check out the Technical details and Model weights. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Liquid AI’s New LFM2-24B-A2B Hybrid Architecture Blends Attention with Convolutions to Solve the Scaling Bottlenecks of Modern LLMs appeared first on MarkTechPost.

Liquid AI’s New LFM2-24B-A2B Hybrid Architecture Blends Attention with Convolutions to Solve the Scaling Bottlenecks of Modern LLMs Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

RAG vs. Context Stuffing: Why selective retrieval is more efficient and reliable than dumping all data into the prompt

Large context windows have dramatically increased how much information modern language models can process in a single prompt. With models capable of handling hundreds of thousands—or even millions—of tokens, it’s easy to assume that Retrieval-Augmented Generation (RAG) is no longer necessary. If you can fit an entire codebase or documentation library into the context window, why build a retrieval pipeline at all? The key distinction is that a context window defines how much the model can see, while RAG determines what the model should see. A large window increases capacity, but it does not improve relevance. RAG filters and selects the most important information before it reaches the model, improving signal-to-noise ratio, efficiency, and reliability. The two approaches solve different problems and are not substitutes for one another. In this article, we compare both strategies directly. Using the OpenAI API, we evaluate Retrieval-Augmented Generation against brute-force context stuffing on the same documentation corpus. We measure token usage, latency, and cost—and demonstrate how burying critical information inside large prompts can affect model performance. The results highlight why large context windows complement RAG rather than replace it. Installing the dependencies Copy CodeCopiedUse a different Browser import os import time import textwrap import numpy as np import tiktoken from openai import OpenAI from getpass import getpass os.environ[“OPENAI_API_KEY”] = getpass(‘Enter OpenAI API Key: ‘) client = OpenAI() We use text-embedding-3-small as the embedding model to convert documents and queries into vector representations for efficient semantic retrieval. For generation and reasoning, we use gpt-4o, with token accounting handled via its corresponding tiktoken encoding to accurately measure context size and cost. Copy CodeCopiedUse a different Browser EMBED_MODEL = “text-embedding-3-small” CHAT_MODEL = “gpt-4o” ENC = tiktoken.encoding_for_model(“gpt-4o”) Creating the document corpus This corpus serves as the retrieval source for our benchmark. In the RAG setup, embeddings are generated for each document and relevant chunks are retrieved based on semantic similarity. In the context-stuffing setup, the entire corpus is injected into the prompt. Because the documents contain specific numeric clauses (e.g., time limits, rate caps, refund windows), they are well-suited for testing retrieval accuracy, signal density, and the “Lost in the Middle” effect under large-context conditions. The corpus consists of 10 structured policy documents totaling approximately 650 tokens, with each document ranging between 54 and 83 tokens. This size keeps the dataset manageable while still reflecting the diversity and density of a realistic enterprise documentation set. Although relatively small, the corpus includes tightly packed numerical clauses, conditional rules, and compliance statements—making it suitable for evaluating retrieval precision, reasoning accuracy, and token efficiency. It provides a controlled environment to compare RAG-based selective retrieval against full context stuffing without introducing external noise. Copy CodeCopiedUse a different Browser def count_tokens(text: str) -> int: return len(ENC.encode(text)) DOCS = [ { “id”: 1, “title”: “Refund Policy”, “content”: ( “Customers may request a full refund within 30 days of purchase. ” “Refunds are processed within 5-7 business days to the original payment method. ” “Digital products are non-refundable once the download link has been accessed. ” “Subscription cancellations stop future charges but do not trigger automatic refunds ” “for the current billing cycle unless the cancellation is made within 48 hours of renewal.” ) }, { “id”: 2, “title”: “Shipping Information”, “content”: ( “Standard shipping takes 5-7 business days. Express shipping delivers in 2-3 business days. ” “Orders over $50 qualify for free standard shipping within the continental US. ” “International shipping is available to 30 countries and takes 10-21 business days. ” “Tracking numbers are emailed within 24 hours of dispatch.” ) }, { “id”: 3, “title”: “Account Security”, “content”: ( “Two-factor authentication (2FA) can be enabled from the Security tab in account settings. ” “Passwords must be at least 12 characters and include one uppercase letter, one number, ” “and one special character. Active sessions expire after 30 days of inactivity. ” “Suspicious login attempts trigger an automatic account lock and a reset email.” ) }, { “id”: 4, “title”: “API Rate Limits”, “content”: ( “Free tier: 100 requests per day, max 10 requests per minute. ” “Pro tier: 10 000 requests per day, max 200 requests per minute. ” “Enterprise tier: unlimited requests, burst up to 1 000 per minute. ” “All responses include X-RateLimit-Remaining and X-RateLimit-Reset headers. ” “Exceeding limits returns HTTP 429 with a Retry-After header.” ) }, { “id”: 5, “title”: “Data Privacy & GDPR”, “content”: ( “All user data is encrypted at rest using AES-256 and in transit using TLS 1.3. ” “We never sell or rent personal data to third parties. ” “The platform is fully GDPR and CCPA compliant. ” “Data deletion requests are processed within 72 hours. ” “Users can export all their data in JSON or CSV format from the Privacy section.” ) }, { “id”: 6, “title”: “Billing & Subscription Cycles”, “content”: ( “Subscriptions renew automatically on the same calendar day each month. ” “Annual plans offer a 20 % discount compared to monthly billing. ” “Invoices are sent via email 3 days before each renewal. ” “Failed payments retry three times over 7 days before the account is downgraded.” ) }, { “id”: 7, “title”: “Supported File Formats”, “content”: ( “Supported upload formats: PDF, DOCX, XLSX, PPTX, PNG, JPG, WebP, MP4, MOV. ” “Maximum individual file size is 100 MB. ” “Batch uploads support up to 50 files simultaneously. ” “Files are virus-scanned on upload and quarantined if threats are detected.” ) }, { “id”: 8, “title”: “Compliance Certifications”, “content”: ( “The platform holds SOC 2 Type II certification, renewed annually. ” “ISO 27001 compliance is maintained with quarterly internal audits. ” “A HIPAA Business Associate Agreement (BAA) is available for healthcare customers on the Enterprise plan. ” “PCI-DSS Level 1 compliance covers all payment processing flows.” ) }, { “id”: 9, “title”: “SLA & Uptime Guarantees”, “content”: ( “Enterprise SLA guarantees 99.9 % monthly uptime (≤ 43 minutes downtime/month). ” “Scheduled maintenance windows occur every Sunday between 02:00-04:00 UTC. ” “Unplanned incidents are communicated via status.example.com within 15 minutes. “

RAG vs. Context Stuffing: Why selective retrieval is more efficient and reliable than dumping all data into the prompt Leggi l'articolo »

We use cookies to improve your experience and performance on our website. You can learn more at Politica sulla privacy and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
it_IT