YouZum

Uncategorized

AI, Committee, Notizie, Uncategorized

The Download: a battery pivot to AI, and rewriting math

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Why this battery company is pivoting to AI  Qichao Hu doesn’t mince words about the state of the battery industry. “Almost every Western battery company has either died or is going to die. It’s kind of the reality,” he says.   Hu is the CEO of SES AI, a Massachusetts-based battery company. It previously developed advanced lithium batteries for major industries, but is now shifting to AI materials discovery. Read our story to find out why.   —Casey Crownhart  This startup wants to change how mathematicians do math  Axiom Math, a California startup, has released a free AI tool with a big ambition: discovering mathematical patterns that could unlock solutions to long-standing problems.  Most of the successes with AI tools have involved finding solutions to existing problems. But that’s not all they could do. There are lots of problems in math that require new ideas nobody has ever had, which could come from spotting patterns that have never been spotted before.   Axiom Math’s new tool aims to find these hidden links. Read the full story to discover their plans—and how AI in general could change mathematics.  —Will Douglas Heaven  Are high gas prices good news for EVs? It’s complicated.  As the conflict in Iran has escalated, fossil-fuel prices have been on a roller-coaster—and some EV owners are celebrating.   They believe the volatility will create an opportunity for electric vehicles to make headway. But even the carless among us should be concerned about a sustained rise in fossil-fuel prices.   To find out why, read the full story.  —Casey Crownhart  This article is from The Spark, our weekly climate newsletter. Sign up to receive it in your inbox every Wednesday.  The must-reads  I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.  1 Meta and YouTube have been fined for designing addictive products They must pay damages of $6 million for harming young people. (Guardian) + The verdicts will reshape legal protections for Big Tech. (WSJ $) + They could also ripple through social media markets worldwide. (Rest of World) + Juries have started taking the lead in the push for child online safety. (NYT)  2 SpaceX aims to file for IPO as soon as this week It’s hoping to raise more than $75 billion. (The Information) + Rocket stocks soared on the report. (BBC)  + But rivals are challenging SpaceX’s dominance. (MIT Technology Review)  3 A new AI safety bill would halt data center construction It was introduced by Bernie Sanders. (Wired) + Nobody wants a data center in their backyard. (MIT Technology Review + One solution: launch them into space. (MIT Technology Review)   4 Meta has laid off 700 employees After raising compensation for top earners. (NYT $)  5 Elon Musk wants a Delaware judge to recuse herself over an emoji She liked a LinkedIn post criticizing him. (CNBC) + The case had ruled Musk misled investors during the Twitter purchase. (Reuters)  6 Reddit will require “fishy” accounts to verify that a human runs them The process aims to combat the deluge of bots. (Ars Technica)  7 Uber and Pony AI aim to launch Europe’s first robotaxi service in Croatia Pony AI is also running trials in Luxembourg, while Uber is testing in London. (The Verge)  8 Google says quantum computers could break all cryptographic security by 2029 It’s set a timeline to secure the quantum era. (Gizmodo) + Quantum computers could soon solve health care problems. (MIT Technology Review)  9 New research shows cloning doesn’t produce perfect copies Clones have lots of extra, potentially dangerous mutations. (New Scientist)  10 The landmark AI Scientist has just completed peer review  It’s billed as the first AI tool built to fully automate the scientific process. (Nature)  Quote of the day  “For years, social media companies have profited from targeting children while concealing their addictive and dangerous design features. Today’s verdict is a referendum—from a jury, to an entire industry.”  —Attorney Rachel Lanier offers her view on yesterday’s fines for Meta and YouTube, the Washington Post reports.   One More Thing  GETTY IMAGES Longevity enthusiasts want to create their own independent state. They’re eyeing Rhode Island.   It’s incredibly difficult and expensive to study innovative ways to slow or reverse aging. In response, longevity enthusiasts have devised an ambitious plan: establish an independent state for life-extension experiments.   They envision a jurisdiction that slashes red tape, encourages self-experimentation with unproven treatments, and eliminates laws that limit how companies develop drugs.   Exactly where their longevity state might emerge is still being worked out—but one appealing location is Rhode Island. Read the full story to learn more about the plans.   —Jessica Hamzelou  We can still have nice things  A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)  + These gleaming photos of ancient insects in amber are time capsules of the dinosaur age. + Paint with pixels across a world map at this unique digital canvas. + Hands have a new shield against hammers: a nail holder that protects your fingers. + This new audio player uses cartridges to give digital music a soul. 

The Download: a battery pivot to AI, and rewriting math Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

The AI Hype Index: AI goes to war

AI is at war. Anthropic and the Pentagon feuded over how to weaponize Anthropic’s AI model Claude; then OpenAI swept the Pentagon off its feet with an “opportunistic and sloppy” deal. Users quit ChatGPT in droves. People marched through London in the biggest protest against AI to date. If you’re keeping score, Anthropic—the company founded to be ethical—is now turbocharging US strikes on Iran.  On the lighter side, AI agents are now going viral online. OpenAI hired the creator of OpenClaw, a popular AI agent. Meta snapped up Moltbook, where AI agents seem to ponder their own existence and invent new religions like Crustafarianism. And on RentAHuman, bots are hiring people to deliver CBD gummies. The future isn’t AI taking your job. It’s AI becoming your boss and finding God.

The AI Hype Index: AI goes to war Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Agentic commerce runs on truth and context

Imagine telling a digital agent, “Use my points and book a family trip to Italy. Keep it within budget, pick hotels we’ve liked before, and handle the details.” Instead of returning a list of links, the agent assembles an itinerary and executes the purchase. That shift, from assistance to execution, is what makes agentic AI different. It also changes the operating speed of commerce. Payment transactions are already clear in milliseconds. The new acceleration is everything before the payment: discovery, comparison, decisioning, authorization, and follow-through across many systems. As humans step out of routine decisions, “good enough” data stops being good enough. In an agent-driven economy, the constraint isn’t speed; it’s trust at machine speed and scale. Automated markets already work because identity, authority, and accountability are built in. As agents transact across businesses, that same clarity is required. Master data management (MDM)—the discipline of creating a single master record—becomes the exchange layer: tracking who an agent represents, what it can do, and where responsibility sits when value moves. Markets don’t fail from automation; they fail from ambiguous ownership. MDM turns autonomous action into legitimate, scalable trust. To make agentic commerce safe and scalable, organizations will need more than better models. They will need a modern data architecture and an authoritative system of context that can instantly recognize, resolve, and distinguish entities. It is the difference between automation that scales and automation that needs constant human correction. The agent is a new participant Digital commerce has long been built on two primary sides: buyers and suppliers/merchants. Agentic commerce adds a third participant that must be treated as a first-class entity: the agent acting on the buyer’s behalf. That sounds simple until you ask the questions every enterprise will face: Who is the individual, across channels and devices, with enough certainty for automation? Who is the agent, and what permissions and limits define what it can do? Who is the merchant or supplier, and are we sure we mean the right one? Who holds liability if the agent acts with permission, but against user intent? The practical risk is confusion. Humans, for example, can infer that “Delta” means the airline when they are booking a flight, not the faucet company. An agent needs deterministic signals. If the system guesses wrong, it either breaks trust or forces a human confirmation step that defeats the promise of speed. Why ‘good enough’ data breaks at machine speed Most organizations have learned to live with imperfect data. Duplicate customer records are tolerable. Incomplete product attributes are annoying. Merchant identities can be reconciled later. Agentic workflows change that tolerance. When an agent takes action without a human checking the output, it needs data that is close to perfect, because it cannot reliably notice when data is ambiguous or wrong the way a person can. The failure modes are predictable, and they show up in places that matter most: Product truth: If the catalog is inconsistent, an agent’s choices will look arbitrary (“the wrong shirt,” “the wrong size,” “the wrong material”), and trust collapses quickly. Payee truth: Agentic commerce expands beyond cards to account-to-account and open-banking-connected experiences, broadening the universe of payees and the need to recognize them accurately in real time. Identity truth: People operate in multiple contexts (work versus personal). Devices shift. A system that cannot distinguish amongst these contexts will either block legitimate activity or approve risky activity, both of which damage adoption. This is why unified enterprise data and entity resolution move from nice to have to operationally required. The more autonomy you want, the more you must invest in modern data foundations that ensure it is safe. Context intelligence: The missing layer When leaders talk about agentic AI, they often focus on model capability: planning, tool use, and reasoning. Those are necessary, but they are not sufficient. Agentic commerce also requires a layer that provides authoritative context at runtime. Think of it as a real-time system of context that can answer instantly and consistently: • Is this the right person?• Is this the right agent, acting within the right permissions?• Is this the right merchant or payee?• What constraints apply right now (budget, policy, risk, loyalty rules, preferred suppliers)? Two design principles matter. First, entity truth must be deterministic enough for automation. Large language models are probabilistic by nature. That is helpful for creating options for writing and drawing. It is risky for deciding where money goes, especially in B2B and finance workflows, where “probably correct” is not acceptable. Second, context must travel at the speed of interaction and remain portable across the entire connected network value chain. Mastercard’s experience optimizing payment flows is instructive: the more services you layer onto a transaction, the more you risk slowing it down. The pattern that scales pre-resolves, curates, and packages the signal so that execution is lightweight. This is also where tokenization is heading. Initiatives like Mastercard’s Agent Pay and Verifiable Intent signal a future in which consumer credentials, agent identities, permissions, and provable user intent are encoded as cryptographically secure artifacts — enabling merchants, issuers and platforms to deterministically verify authorization and execution at machine speed. What leaders should do in the next 12 to 24 months Adoption will not be uniform. Early traction will often depend less on industry and more on the sophistication of an organization’s systems and data discipline. That makes the next two years a window for practical preparation. Five moves stand out. Treat agents as governed identities, not features. Define how agents are onboarded, authenticated, permissioned, monitored, and retired. Prioritize entity resolution where the cost of being wrong is highest. Start with payees, suppliers, employee-versus-personal identity, and high-volume product categories. Build a reusable context service that every workflow and agent can call. Do not force each system to reconstruct identity and relationships from scratch. Precompute and compress signals. Resolve and curate context upstream so that runtime decisioning stays fast and predictable. Expand autonomy only as trust is earned. Build a governance framework to address disputes, keep humans in the loop

Agentic commerce runs on truth and context Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

The Download: reawakening frozen brains, and the AI Hype Index returns

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. This scientist rewarmed and studied pieces of his friend’s cryopreserved brain  L. Stephen Coles’s brain sits in a vat at a storage facility in Arizona. It has been held there at a temperature of around −146 degrees °C for over a decade, largely undisturbed. Before he died in 2014, Coles had the brain frozen with an ambitious goal in mind: reanimation.  His friend, cryobiologist Greg Fahy, believes it could be revived one day. But other experts are less optimistic.   Still, Fahy’s research could lead to new ways to study the brain. And using cryopreservation for organ transplantation is becoming a viable reality.   Read the full story to find out what the future holds for the technology.  —Jessica Hamzelou  The AI Hype Index  Separating AI reality from hyped-up fiction isn’t always easy. That’s why we’ve created the AI Hype Index—a simple, at-a-glance summary of everything you need to know about the state of the industry. Take a look at this month’s edition.   MIT Technology Review Narrated: how Pokémon Go is giving delivery robots an inch-perfect view of the world   Pokémon Go was the world’s first augmented-reality megahit. Released in 2016 by Niantic, the AR twist on the juggernaut Pokémon franchise fast became a global phenomenon. “500 million people installed that app in 60 days,” says Brian McClendon, CTO at Niantic Spatial, an AI company that Niantic spun out last year.   Now Niantic Spatial is using that vast trove of crowdsourced data to build a kind of world model—a buzzy new technology that grounds the smarts of LLMs in real environments. The firm wants to use it to help robots navigate more precisely.  —Will Douglas Heaven  This is our latest story to be turned into an MIT Technology Review Narrated podcast, which we’re publishing each week on Spotify and Apple Podcasts. Just navigate to MIT Technology Review Narrated on either platform, and follow us to get all our new content as it’s released.  The next era of space exploration  Our footprint in the solar system is rapidly expanding. Programs to build permanent Moon bases and find life on Mars have transitioned from science fiction to active space agency missions. The scientists behind them will not only shed new light on the cosmos, but also reveal where humanity is headed.  To examine what the future holds in store, MIT Technology Review features editor Amanda Silverman will sit down today with award-winning science journalist and author Robin George Andrews for an exclusive subscriber-only Roundtable conversation about “The Next Era of Space Exploration.” Register here to join the session at 16:00 GMT / 12:00 PM ET / 9:00 AM PT.  The must-reads  I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.  1 OpenAI is shutting down AI video generator Sora  The app attracted at least as much controversy as acclaim. (CNBC) + Closing it means saying goodbye to $1 billion from Disney. (BBC) + OpenAI is cutting back on side projects ahead of an expected IPO. (WSJ $) + But it’s focusing its efforts on building a fully automated researcher. (MIT Technology Review)  2 A judge suspects the Pentagon is illegally punishing Anthropic She labelled the DoD’s ban “troubling.” (Bloomberg) + Anthropic and the Pentagon are facing off in court. (Guardian) + The DoD wants AI companies to train on classified data. (MIT Technology Review)  3 Meta has been ordered to pay $375 million for endangering children online Prosecutors said the company knew it put children at risk. (Engadget) + Meta is offering its top talent stock options as incentives for its AI push. (CNBC)  4 Arm will sell its own computer chips for the first time It’s aimed at data centers that run AI tasks. (NYT $) + Arm stock jumped 13% on the news. (CNBC)  5 Manus’s founders have been barred from leaving China following Meta’s takeover Beijing is reviewing the $2 billion acquisition of the AI startup. (FT $)  6 Baltimore has sued xAI over Grok’s fake nude images  The chatbot allegedly violated consumer protections. (Guardian) + There’s a big market for pornographic deepfakes of real women. (MIT Technology Review)  7 NASA plans to send a nuclear-powered spacecraft to Mars in 2028 It’ll take a payload of Ingenuity-class helicopters to the Red Planet. (NYT $) + NASA also wants to put a $20 billion base on the Moon. (The Verge)  8 A company is secretly turning Zoom meetings into AI-generated podcasts WebinarTV turns the calls into content without telling anyone. (404 Media)  9 Iranian volunteers have built their own missile warning map It fills the gap left by Iran’s lack of a public emergency alert tool. (Wired $) + Here’s where OpenAI’s tech could show up in Iran. (MIT Technology Review)  10 A nonprofit is sending basic income payments to AI-impacted workers It’s starting by giving 25-50 people $1,000 per month. (Gizmodo)  Quote of the day  “I am first and foremost a scientist. My goal is to understand nature. But doing science is, sort of, like reading the mind of God.”  —DeepMind CEO Demis Hassabis shares his approach to AI strategy with the FT.  One More Thing  EVA REDAMONTI Inside the hunt for the most dangerous asteroid ever   As asteroid 2024 YR4 hurtled toward Earth, astronomers determined that this massive rock posed a higher risk of impact than any object of its size in recorded history. Then, just as quickly as history was made, experts declared that the danger had passed.  This is the inside story of the network of global scientists who found, followed, planned for, and finally dismissed the most dangerous asteroid ever found—all under the tightest of timelines and with the highest of stakes. Find out how they did it.  —Robin George Andrews  We can still have nice things  A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)  + Soothe subscription fatigue with this simple cancellation tool. + Takashi Murakami’s reimagined Monets are pop-art magic. + Jump into a rabbit hole with this app that visualizes links between Wikipedia pages. + This playful lynx that snatched the top prize in a photo competition is a delight. 

The Download: reawakening frozen brains, and the AI Hype Index returns Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

This startup wants to change how mathematicians do math

Axiom Math, a startup based in Palo Alto, California, has released a free new AI tool for mathematicians, designed to discover mathematical patterns that could unlock solutions to long-standing problems. The tool, called Axplorer, is a redesign of an existing one called PatternBoost that François Charton, now a research scientist at Axiom, co-developed in 2024 when he was at Meta. PatternBoost ran on a supercomputer; Axplorer runs on a Mac Pro. The aim is to put the power of PatternBoost, which was used to crack a hard math puzzle known as the Turán four-cycles problem, in the hands of anyone who can install Axplorer on their own computer. Last year, the US Defense Advanced Research Projects Agency set up a new initiative called expMath—short for Exponentiating Mathematics—to encourage mathematicians to develop and use AI tools. Axiom sees itself as part of that drive. Breakthroughs in math have enormous knock-on effects across technology, says Charton. In particular, new math is crucial for advances in computer science, from building next-generation AI to improving internet security. Most of the successes with AI tools have involved finding solutions to existing problems. But finding solutions is not all that mathematicians do, says Axiom Math founder and CEO Carina Hong. Math is exploratory and experimental, she says.  MIT Technology Review met with Charton and Hong last week for an exclusive video chat about their new tool and how AI in general could change mathematics.  Math by chatbot In the last few months, a number of mathematicians have used LLMs, such as OpenAI’s GPT-5, to find solutions to unsolved problems, especially ones set by the 20th-century mathematician Paul Erdős, who left behind hundreds of puzzles when he died. But Charton is dismissive of those successes. “There are tons of problems that are open because nobody looked at them, and it’s easy to find a few gems you can solve,” he says. He’s set his sights on tougher challenges—“the big problems that have been very, very well studied and famous people have worked on them.” Last year, Axiom Math used another of its tools, called AxiomProver, to find solutions to four such problems in mathematics.    The Turán four-cycles problem that PatternBoost cracked is another big problem, says Charton. (The problem is an important one in graph theory, a branch of math that’s used to analyze complex networks such as social media connections, supply chains, and search engine rankings. Imagine a page covered in dots. The puzzle involves figuring out how to draw lines between as many of the dots as possible without creating loops that connect four dots in a row.) “LLMs are extremely good if what you want to do is derivative of something that has already been done,” says Charton. “This is not surprising—LLMs are pretrained on all the data that there is. But you could say that LLMs are conservative. They try to reuse things that exist.” However, there are lots of problems in math that require new ideas, insights that nobody has ever had. Sometimes those insights come from spotting patterns that hadn’t been spotted before. Such discoveries can open up whole new branches of mathematics. PatternBoost was designed to help mathematicians find new patterns. Give the tool an example and it generates others like it. You select the ones that seem interesting and feed them back in. The tool then generates more like those, and so on.   It’s a similar idea to Google DeepMind’s AlphaEvolve, a system that uses an LLM to come up with novel solutions to a problem. AlphaEvolve keeps the best suggestions and asks the LLM to improve on them. Special access Researchers have already used both AlphaEvolve and PatternBoost to discover new solutions to long-standing math problems. The trouble is that those tools run on large clusters of GPUs and are not available to most mathematicians. Mathematicians are excited about AlphaEvolve, says Charton. “But it’s closed—you need to have access to it. You have to go and ask the DeepMind guy to type in your problem for you.” And when Charton solved the Turán problem with PatternBoost, he was still at Meta. “I had literally thousands, sometimes tens of thousands, of machines I could run it on,” he says. “It ran for three weeks. It was embarrassing brute force.” Axplorer is far faster and far more efficient, according to the team at Axiom Math. Charton says it took Axplorer just 2.5 hours to match PatternBoost’s Turán result. And it runs on a single machine. Geordie Williamson, a mathematician at the University of Sydney, who worked on PatternBoost with Charton, has not yet tried Axplorer. But he is curious to see what mathematicians do with it. (Williamson still occasionally collaborates with Charton on academic projects but says he is not otherwise connected to Axiom Math.) Williamson says Axiom Math has made several improvements to PatternBoost that (in theory) make Axplorer applicable to a wider range of mathematical problems. “It remains to be seen how significant these improvements are,” he says. “We are in a strange time at the moment, where lots of companies have tools that they’d like us to use,” Williamson adds. “I would say mathematicians are somewhat overwhelmed by the possibilities. It is unclear to me what impact having another such tool will be.” Hong admits that there are a lot of AI tools being pitched at mathematicians right now. Some also require mathematicians to train their own neural networks. That’s a turnoff, says Hong, who is a mathematician herself. Instead, Axplorer will walk you through what you want to do step by step, she says. The code for Axplorer is open source and available via GitHub. Hong hopes that students and researchers will use the tool to generate sample solutions and counterexamples to problems they’re working on, speeding up mathematical discovery. Williamson welcomes new tools and says he uses LLMs a lot. But he doesn’t think mathematicians should throw out the whiteboards just yet. “In my biased opinion, PatternBoost is a lovely idea, but it is

This startup wants to change how mathematicians do math Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

arXiv:2601.03641v3 Announce Type: replace Abstract: Large Language Model (LLM)-based agents significantly extend the utility of LLMs by interacting with dynamic environments. However, enabling agents to continually learn new tasks without catastrophic forgetting remains a critical challenge, known as the stability-plasticity dilemma. In this work, we argue that this dilemma fundamentally arises from the failure to explicitly distinguish between common knowledge shared across tasks and conflicting knowledge introduced by task-specific interference. To address this, we propose Agent-Dice, a parameter fusion framework based on directional consensus evaluation. Concretely, Agent-Dice disentangles knowledge updates through a two-stage process: geometric consensus filtering to prune conflicting gradients, and curvature-based importance weighting to amplify shared semantics. We provide a rigorous theoretical analysis that establishes the validity of the proposed fusion scheme and offers insight into the origins of the stability-plasticity dilemma. Extensive experiments on GUI agents and tool-use agent domains demonstrate that Agent-Dice exhibits outstanding continual learning performance with minimal computational overhead and parameter updates. The codes are available at https://github.com/Wuzheng02/Agent-Dice.

Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

NDT: Non-Differential Transformer and Its Application to Sentiment Analysis

arXiv:2603.20704v1 Announce Type: cross Abstract: From customer feedback to social media, understanding human sentiment in text is central to how machines can interact meaningfully with people. However, despite notable progress, accurately capturing sentiment remains a challenging task, which continues to motivate further research in this area. To this end, we introduce Non-Differential Transformer (NDT). It is inspired by (but in contrast to) the state-of-the-art Differential Transformer (DT) model. While standard Transformers can struggle with irrelevant context, the sota DT model uses attention map subtraction, potentially for noise cancellation. We explore an alternative motivation, hypothesizing that benefits may arise from enabling different attention components to specialize on distinct concepts within the text, similar to multiplexing information channels or mixture models, rather than primarily canceling noise via subtraction. Guided by this concept-multiplexing (ConPlex) view, the specific architecture presented in this paper employs a purely additive strategy. It uses only positive weights, learned during training, to ensure constructive combination of these specialized attention perspectives. This design choice explores positive only integration, though our broader framework also shows promise with less constrained linear combinations involving both positive and negative weights. Our model computes attention via this positively weighted sum of multiple distinct attention maps. This allows the model to constructively integrate diverse signals and potentially capture more complex contextual relationships. Competitive performance is achieved by the proposed model for Sentiment Analysis while tested on multiple datasets. We conclude by presenting our results, challenges and future research agenda in this important area of research.

NDT: Non-Differential Transformer and Its Application to Sentiment Analysis Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Yann LeCun’s New LeWorldModel (LeWM) Research Targets JEPA Collapse in Pixel-Based Predictive World Modeling

World Models (WMs) are a central framework for developing agents that reason and plan in a compact latent space. However, training these models directly from pixel data often leads to ‘representation collapse,’ where the model produces redundant embeddings to trivially satisfy prediction objectives. Current approaches attempt to prevent this by relying on complex heuristics: they utilize stop-gradient updates, exponential moving averages (EMA), and frozen pre-trained encoders. A team of researchers including Yann LeCun and many others (Mila & Université de Montréal, New York University, Samsung SAIL and Brown University) introduced LeWorldModel (LeWM), the first JEPA (Joint-Embedding Predictive Architecture) that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings Technical Architecture and Objective LeWM consists of two primary components learned jointly: an Encoder and a Predictor. Encoder ((zt=encθ (ot)): Maps a raw pixel observation into a compact, low-dimensional latent representation. The implementation uses a ViT-Tiny architecture (~5M parameters). Predictor (Žt+1=predθ(zt, at)): A transformer (~10M parameters) that models environment dynamics by predicting future latent states conditioned on actions. The model is optimized using a streamlined objective function consisting of only two loss terms: $$mathcal{L}_{LeWM} triangleq mathcal{L}_{pred} + lambda SIGReg(Z)$$ The prediction loss (Lpred) computes the mean-squared error (MSE) between the predicted and actual consecutive embeddings. The SIGReg (Sketched-Isotropic-Gaussian Regularizer) is the anti-collapse term that enforces feature diversity. As per the research paper, applying a dropout rate of 0.1 in the predictor and a specific projection step (1-layer MLP with Batch Normalization) after the encoder are critical for stability and downstream performance. Efficiency via SIGReg and Sparse Tokenization Assessing normality in high-dimensional latent spaces is a major scaling challenge. LeWM addresses this using SIGReg, which leverages the Cramér-Wold theorem: a multivariate distribution matches a target (isotropic Gaussian) if all its one-dimensional projections match that target. SIGReg projects latent embeddings onto M random directions and applies the Epps-Pulley test statistic to each resulting one-dimensional projection. Because the regularization weight λ is the only effective hyperparameter to tune, researchers can optimize it using a bisection search with O(log n) complexity, a significant improvement over the polynomial-time search (O(n6)) required by previous models like PLDM. Speed Benchmarks In the reported setup, LeWM demonstrates high computational efficiency: Token Efficiency: LeWM encodes observations using ~200× fewer tokens than DINO-WM. Planning Speed: LeWM achieves planning up to 48× faster than DINO-WM (0.98s vs 47s per planning cycle). Latent Space Properties and Physical Understanding LeWM’s latent space supports probing of physical quantities and detection of physically implausible events. Violation-of-Expectation (VoE) Using a VoE framework, the model was evaluated on its ability to detect ‘surprise’. It assigned higher surprise to physical perturbations such as teleportation; visual perturbations produced weaker effects, and cube color changes in OGBench-Cube were not significant. Emergent Path Straightening LeWM exhibits Temporal Latent Path Straightening, where latent trajectories naturally become smoother and more linear over the course of training. Notably, LeWM achieves higher temporal straightness than PLDM despite having no explicit regularizer encouraging this behavior. Feature LeWorldModel (LeWM) PLDM DINO-WM Dreamer / TD-MPC Training Paradigm Stable End-to-End End-to-End Frozen Foundation Encoder Task-Specific Input Type Raw Pixels Raw Pixels Pixels (DINOv2 features) Rewards / Privileged State Loss Terms 2 (Prediction + SIGReg) 7 (VICReg-based) 1 (MSE on latents) Multiple (Task-specific) Tunable Hyperparams 1 (Effective weight λ) 6 N/A (Fixed by pre-training) Many (Task-dependent) Planning Speed Up to 48x Faster Fast (Compact latents) Slow (~50x slower than LeWM) Varies (often slow generation) Anti-Collapse Provable (Gaussian prior) Under-specified / Unstable Bounded by pre-training Heuristic (e.g., reconstruction) Requirement Task-Agnostic / Reward-Free Task-Agnostic / Reward-Free Frozen Pre-trained Encoder Task Signals / Rewards Key Takeaways Stable End-to-End Learning: LeWM is the first Joint-Embedding Predictive Architecture (JEPA) that trains stably end-to-end from raw pixels without needing ‘hand-holding’ heuristics like stop-gradients, exponential moving averages (EMA), or frozen pre-trained encoders. A Radical Two-Term Objective: The training process is simplified into just two loss terms—a next-embedding prediction loss and the SIGReg regularizer—reducing the number of tunable hyperparameters from six to one compared to existing end-to-end alternatives. Built for Real-Time Speed: By representing observations with approximately 200× fewer tokens than foundation-model-based counterparts, LeWM plans up to 48× faster, completing full trajectory optimizations in under one second. Provable Anti-Collapse: To prevent the model from learning ‘garbage’ redundant representations, it uses the SIGReg regularizer; this utilizes the Cramér-Wold theorem to ensure high-dimensional latent embeddings stay diverse and Gaussian-distributed. Intrinsic Physical Logic: The model doesn’t just predict data; it captures meaningful physical structure in its latent space, allowing it to accurately probe physical quantities and detect ‘impossible’ events like object teleportation through a violation-of-expectation framework. Check out the Paper, Website and Repo. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Yann LeCun’s New LeWorldModel (LeWM) Research Targets JEPA Collapse in Pixel-Based Predictive World Modeling appeared first on MarkTechPost.

Yann LeCun’s New LeWorldModel (LeWM) Research Targets JEPA Collapse in Pixel-Based Predictive World Modeling Leggi l'articolo »

We use cookies to improve your experience and performance on our website. You can learn more at Politica sulla privacy and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
it_IT