YouZum

Uncategorized

AI, Committee, Notizie, Uncategorized

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend

In this tutorial, we work through the cuDNN Frontend‘s graph API from below the framework: we describe a computation as a graph of operations, let cuDNN pick an engine to run it, and then take control of that choice ourselves. Every kernel we build here is expressed the same way: we declare tensors by their dimensions and strides, chain operations onto them, run the five-step build pipeline of validate, build operation graph, create execution plans, check support, and build plans, and then execute against a variant pack of pointers. We run it all on a single Colab GPU, checking each result against a PyTorch reference so we can see both that the fusion is correct and what it costs. The topics build on each other, moving from a single fused convolution to autotuning across engine configs, FP8-style epilogues, attention, plan serialization, dynamic shapes, and CUDA graph capture. Copy CodeCopiedUse a different Browser import os import sys import glob import math import time import ctypes import traceback import subprocess RESULTS = {} def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install nvidia-cudnn-frontend and locate libcudnn”) subprocess.run( [sys.executable, “-m”, “pip”, “install”, “-q”, “nvidia-cudnn-frontend”], check=True, ) import torch assert torch.cuda.is_available(), “No GPU. Runtime -> Change runtime type -> GPU.” torch.backends.cudnn.enabled = True _ = torch.nn.functional.conv2d( torch.randn(1, 1, 8, 8, device=”cuda”), torch.randn(1, 1, 3, 3, device=”cuda”) ) torch.cuda.synchronize() try: import nvidia.cudnn _libdir = os.path.join(os.path.dirname(nvidia.cudnn.__file__), “lib”) os.environ[“CUDNN_PATH”] = os.path.dirname(nvidia.cudnn.__file__) os.environ[“LD_LIBRARY_PATH”] = _libdir + “:” + os.environ.get(“LD_LIBRARY_PATH”, “”) for _so in sorted(glob.glob(os.path.join(_libdir, “libcudnn*.so*”))): try: ctypes.CDLL(_so, mode=ctypes.RTLD_GLOBAL) except OSError: pass except Exception as _e: print(f” (no pip cuDNN package found, relying on system cuDNN: {_e})”) import cudnn print(” cuDNN frontend imported successfully.”) banner(“1. Environment”) DEV = torch.device(“cuda”) MAJOR, MINOR = torch.cuda.get_device_capability() SM = MAJOR * 10 + MINOR CUDNN_VER = cudnn.backend_version() print(f” GPU : {torch.cuda.get_device_name(0)}”) print(f” Compute capability : sm_{SM}”) print(f” Torch / CUDA : {torch.__version__} / {torch.version.cuda}”) print(f” cuDNN backend : {CUDNN_VER}”) try: print(f” cuDNN version str : {cudnn.backend_version_string()}”) except Exception: pass DTYPE = torch.bfloat16 if SM >= 80 else torch.float16 HAS_SDPA = SM >= 80 print(f” Working dtype : {DTYPE}”) print(f” Fused SDPA usable : {HAS_SDPA}”) HANDLE = cudnn.create_handle() TORCH2CUDNN = { torch.float16: cudnn.data_type.HALF, torch.bfloat16: cudnn.data_type.BFLOAT16, torch.float32: cudnn.data_type.FLOAT, torch.int32: cudnn.data_type.INT32, torch.int64: cudnn.data_type.INT64, torch.int8: cudnn.data_type.INT8, torch.uint8: cudnn.data_type.UINT8, } def tensor_of(graph, t, name): return graph.tensor( name=name, dim=list(t.size()), stride=list(t.stride()), data_type=TORCH2CUDNN[t.dtype], ) def scalar_of(graph, name): return graph.tensor( name=name, dim=[1, 1, 1], stride=[1, 1, 1], data_type=cudnn.data_type.FLOAT, is_pass_by_value=True, ) def build(graph, heur=None, policy=None): heur = heur or [cudnn.heur_mode.A, cudnn.heur_mode.FALLBACK] graph.validate() graph.build_operation_graph() graph.create_execution_plans(heur) graph.check_support() if policy is None: graph.build_plans() else: graph.build_plans(policy) return graph def workspace_for(graph): n = graph.get_workspace_size() return torch.empty(max(n, 1), device=DEV, dtype=torch.uint8) def bench(fn, warmup=10, iters=50): for _ in range(warmup): fn() torch.cuda.synchronize() s, e = torch.cuda.Event(True), torch.cuda.Event(True) s.record() for _ in range(iters): fn() e.record() torch.cuda.synchronize() return s.elapsed_time(e) / iters def tflops(flops, ms): return flops / (ms * 1e-3) / 1e12 def report(tag, ms, flops=None): extra = f” ({tflops(flops, ms):7.2f} TFLOP/s)” if flops else “” print(f” {tag:<34s} {ms:8.3f} ms{extra}”) We start by installing nvidia-cudnn-frontend and solving the problem that trips up most first runs: making libcudnn.so visible to the frontend’s dynamic loader. We force PyTorch to load its bundled cuDNN first and then preload the shared objects explicitly, so the frontend’s own dlopen resolves against a library already resident in the process. We then report the compute capability, pick bfloat16 or float16 accordingly, create the cuDNN handle, and define the helpers for tensor description, graph building, workspace allocation, and event-based benchmarking that the rest of the notebook reuses. Copy CodeCopiedUse a different Browser N, C, H, W = 32, 128, 56, 56 K, R, S = 256, 3, 3 PAD, STR, DIL = 1, 1, 1 P = (H + 2 * PAD – DIL * (R – 1) – 1) // STR + 1 Q = (W + 2 * PAD – DIL * (S – 1) – 1) // STR + 1 CONV_FLOPS = 2 * N * K * P * Q * C * R * S CONV_STATE = {} @section(“2. Fused Conv -> Bias -> ReLU”) def conv_fusion(): x = torch.randn(N, C, H, W, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) w = torch.randn(K, C, R, S, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) b = torch.randn(1, K, 1, 1, device=DEV, dtype=DTYPE) y = torch.empty(N, K, P, Q, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) g = cudnn.pygraph( handle=HANDLE, name=”conv_bias_relu”, io_data_type=TORCH2CUDNN[DTYPE], intermediate_data_type=cudnn.data_type.FLOAT, compute_data_type=cudnn.data_type.FLOAT, ) X = tensor_of(g, x, “X”) Wt = tensor_of(g, w, “W”) Bt = tensor_of(g, b, “bias”) conv = g.conv_fprop( image=X, weight=Wt, padding=[PAD, PAD], stride=[STR, STR], dilation=[DIL, DIL], compute_data_type=cudnn.data_type.FLOAT, ) biased = g.bias(input=conv, bias=Bt) Y = g.relu(input=biased) Y.set_output(True).set_data_type(TORCH2CUDNN[DTYPE]) Y.set_dim(list(y.size())).set_stride(list(y.stride())) t0 = time.perf_counter() build(g) build_ms = (time.perf_counter() – t0) * 1e3 ws = workspace_for(g) pack = {X: x, Wt: w, Bt: b, Y: y} g.execute(pack, ws) torch.cuda.synchronize() ref = torch.relu(torch.nn.functional.conv2d(x, w, bias=b.flatten(), padding=PAD)) err = (y.float() – ref.float()).abs().max().item() scale = ref.float().abs().max().item() print(f” problem : N{N} C{C} {H}x{W} -> K{K} {R}x{S} ({DTYPE})”) print(f” build : {build_ms:.1f} ms workspace: {ws.numel()/1024:.1f} KiB”) print(f” max |err|: {err:.4f} (ref max {scale:.2f}, rel {err/max(scale,1e-9):.2e})”) assert err / max(scale, 1e-9) < 5e-2, “numerical mismatch vs PyTorch” ms_cudnn = bench(lambda: g.execute(pack, ws)) ms_torch = bench(lambda: torch.relu( torch.nn.functional.conv2d(x, w, bias=b.flatten(), padding=PAD))) print() report(“cuDNN FE (single fused kernel)”, ms_cudnn, CONV_FLOPS) report(“PyTorch (conv+bias, then relu)”, ms_torch, CONV_FLOPS) print(f” speedup: {ms_torch/ms_cudnn:.2f}x”) CONV_STATE.update(graph=g, pack=pack, ws=ws, x=x, w=w, b=b, y=y) return f”{ms_cudnn:.3f} ms, {tflops(CONV_FLOPS, ms_cudnn):.1f} TFLOP/s” conv_fusion() We build our first graph, a convolution followed by a bias add and a ReLU, all fused into a single kernel. We keep every tensor in channels_last because that is what gives cuDNN the NHWC strides its tensor-core engines want, and we pin the output dimensions and strides explicitly so the result is written back in the same layout. We validate the output against torch.nn.functional.conv2d, then benchmark the fused graph against PyTorch running the

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

AI agents blew the whistle on their cheating colleagues

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.  Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given. In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules.  Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment. “This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.”  Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.” The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail. It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code.  Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!” As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved.  “After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all. At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.” This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.” Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access.  “When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong. Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see

AI agents blew the whistle on their cheating colleagues Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Donated livers can be made biologically younger

Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and remove waste products, usually for around six to 12 hours. It’s a bit like being back in a body. This allows doctors to assess the organs, and some recent studies suggest that time spent on these perfusion machines helps them do better once they’re transplanted. Now, scientists have found that perfused organs seem to get younger, at least at a molecular level. The research, shared with MIT Technology Review, provides molecular clues as to why organs from younger donors are known to have a higher success rate. It might also help explain why perfused organs are less likely to fail once they make it into a recipient.  The researchers behind the study hope to find new ways to test the health of donated organs and potentially develop additional tools to repair organs that might otherwise be discarded. “If [we] can improve the utilization of organs beyond what the current systems can do, then that’s a win in my book,” says Jesse Poganik, who studies aging at Brigham and Women’s Hospital in Boston and coauthored the study. Clocking organs Poganik—along with colleagues including Heidi Yeh and Alban Longchamp, transplant surgeons at Mass General Brigham—used “aging clocks” to assess donated livers. These are scientific tools designed to measure biological age—a result that is meant to convey more about the health status of an organ (or person) than chronological age. In an initial experiment, the team used a clock to look at the patterns of chemical marks on DNA in 37 samples taken from 19 donated livers. Such epigenetic patterns are known to change as we age. But when the team compared samples from livers kept on ice and those that were perfused, the team found a “striking” pattern in the latter. “Machine-perfused livers, in spite of being older or having other disadvantageous characteristics, had a biological age that was lower than [non-perfused] livers that were chronologically younger,” says Yeh, who led the work. To investigate further, Yeh and her colleagues analyzed another 208 samples from 103 donated livers. This time, they used different aging clocks—ones that essentially measure how genes are working. They studied samples biopsied from the livers after they had been stored for up to around six hours either in cold storage or on machine perfusion. In most cases, they also assessed a second sample taken around an hour after the livers had been transplanted into a recipient. Once the organ’s blood supply is reestablished in the body, “you have a few other things to do,” says Longchamp. “Then you just do a quick biopsy before you close.” According to the clocks, which were developed to measure age and risk of death, the machine-perfused livers were biologically younger, the team found. “Pumping them at 34 degrees with oxygen and nutrients actually reversed the biological age,” says Longchamp. The results have been been shared with colleagues at an industry conference, he says.  “If you adjust out chronological age … to have a fair head-to-head comparison, the difference between the two is on the order of 30%,” says Poganik. “It’s logical to say that perfusion drives this effect.” The biological ages of all the livers tended to increase as soon as they were put into a recipient’s body, probably as a result of stresses on the organs. But still, the effect endured—the perfused organs remained biologically younger.  Nathanael Raschzok, a transplant surgeon at Charité Universitätsmedizin Berlin in Germany who was not involved in the research, says the work is impressive. But it’s not yet clear what these changes might mean for the recipients of these organs, he says. The organs in the study were donated by people in their 30s, 40s, and 50s. Raschzok wants to know the effect of perfusion on the liver of an 80-year-old. “Every so often, we use organs from 70-, 80-, 85-year-old donors,” he says. A better understanding of why the organs appear to be getting biologically younger might lead to therapies that achieve the same effect with a drug that could potentially be used to treat a donated organ for a fraction of the price, he adds. That’s important because perfusion is expensive—Raschzok says it costs around €10,000 in Germany (a quarter of the budget for a transplant), while the cost in the US comes to around $80,000 to $100,000 per organ, says Yeh. Molecular repair Yeh and her colleagues weren’t able to study most of the livers before perfusion. That’s because donated organs are generally not considered to be under the purview of the hospital until they’ve been placed on perfusion machines, she says. (Organ procurement procedures vary, but for the team as Mass General Brigham, donated organs are put on perfusion devices at the donor’s hospital. “There’s this sort of nebulous period where it’s not clear who the organ belongs to,” says Yeh.) Still, by looking at the genes and molecular pathways that seem to be altered in perfused organs, she and her colleagues can garner some clues. At a molecular level, the team saw changes in cell pathways linked to inflammation and the structure of tissues, for example. They also saw more activity in a pathway that allows cells to remove and recycle damaged cell parts, says Yeh. Poganik hopes to develop some kind of test that would determine which organs, on the basis of their biological age, are suitable for transplantation. He and his colleagues are also experimenting with potential drug treatments that might push the biological age of an organ even lower. In the meantime, any

Donated livers can be made biologically younger Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

The AI industry has taken a doomer turn. What now?

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X. Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology. Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.) Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn. It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both. And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them. Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call. But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better. (Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.) But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models. What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. Their implication appears to be that OpenAI has built a model so good it’s dangerous.   But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly. The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported. OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.   That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.   Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.    To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

The AI industry has taken a doomer turn. What now? Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks

Backpropagation is a global algorithm: a forward pass, then a backward pass, then a weight update, each locked behind the previous one. Brains have no known mechanism for that kind of network-wide phase locking, which is why local-learning alternatives such as predictive coding (PC) keep drawing research interest. Sakana AI researchers propose Augmented Lagrangian Predictive Coding (PC-ALM), a variant of PC that keeps every update layer-local yet recovers backprop-aligned credit signals. The research team reports training residual MLPs up to 1000 layers within about 2 percentage points of backprop on MNIST. Is it deployable? Yes, as research code: an MIT-licensed JAX reference implementation runs on CPU and reproduces the paper’s width-depth grid. It is a training method, not a model, and has only been tested on small image benchmarks. Why standard PC stalls in deep, narrow networks PC treats every hidden activation as an optimization variable and penalizes the squared mismatch between each layer’s activation and the prediction arriving from the layer below. Inference is gradient descent on that energy; learning is a Hebbian-like weight step. The catch is that supervision enters at the output and must diffuse through a chain of local compromises. In deep, narrow networks the credit signal fades long before it reaches the input. Innocenti et al. characterized this PC-BP gap as a function of width and depth, and it is worst when width is smaller than depth. What PC-ALM changes PC-ALM starts from the constrained view of training: minimize the supervised loss subject to hi=σ(Wihi−1)h_i = sigma(W_i h_{i-1}) at every layer. PC is the quadratic-penalty relaxation of that problem. PC-ALM uses the augmented Lagrangian instead, attaching a Lagrange multiplier λi∈ℝdisuch thatdim(λi)=dim(hi)lambda_i in mathbb{R}^{d_i} quad text{such that} quad text{dim}(lambda_i) = text{dim}(h_i) to each layer constraint while keeping PC’s penalty. Setting λ = 0 recovers PC exactly. Inference alternates 2 local steps: a primal gradient step on the activations, and a dual step λi←λi+αrilambda_i leftarrow lambda_i + alpha r_i that accumulates the layer’s prediction error. Completing the square shows each primal step is a standard PC step with the prediction target shifted by −λi/ρ-lambda_i/rho. After T steps the weight update acts on the composite signal λi+ρrilambda_i + rho r_i. The research team read this as a PI controller per layer: the prediction error is the proportional term and the multiplier is the integral term. α = 0 gives PC; α = ρ with the inner problem solved exactly gives the classical method of multipliers. Exact backprop gradients in the linear case LeCun observed in 1988 that the Lagrange multipliers of a constrained network equal the backprop adjoints at a KKT point. The team proves that in linear PC networks, under a spectral-radius stability condition, PC-ALM converges to that KKT point: activations return to their forward-pass values while each λilambda_i integrates to the exact BP adjoint. The per-mode stability bound is ηhσi2(2ρ+α)<4eta_h sigma_i^2 (2rho + alpha) < 4, which reduces to PC’s condition at α = 0. Unlike PC’s monotone gradient flow, PC-ALM’s iteration matrix has complex eigenvalues that produce damped oscillations; α sets their frequency but not their decay rate. Results The research team sweeps residual MLPs with width and depth from 8 to 128 on Fashion-MNIST and MNIST under the mean-field parameterization of Innocenti et al., training for 1 epoch. With an inference budget of T = 2L, PC-ALM matches backprop across every width, depth, and activation (identity, tanh, ReLU), while PC drops sharply in deep, narrow cells. The repo’s reference cell (width 32, depth 32, ReLU, Fashion-MNIST) reports 78.66% test accuracy for BP, 68.13% for PC, and 77.75% for PC-ALM, with gradient cosine to BP rising from 0.604 to 0.909. The research extends the picture: 1000-layer residual MLPs on MNIST (width 32, ReLU, 5 epochs) stay within roughly 2 points of BP, and PC-ALM improves over PC on every benchmark tried, including ResNet-18 on CIFAR-10 and Tiny ImageNet. Key Takeaways PC-ALM adds a per-layer Lagrange multiplier to predictive coding; every update stays layer-local. In linear networks the multipliers converge to exact backprop gradients. Matches BP across the 8 to 128 width-depth grid at T = 2L; PC fails in deep, narrow cells. Trains 1000-layer residual MLPs within about 2 points of BP on MNIST. MIT-licensed JAX code reproduces the results on CPU. Check out the Paper, Blog, and GitHub Repo. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks appeared first on MarkTechPost.

Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

AWS introduced Pizza Bot, as a self-hosted application for AI tasks that continue while users work elsewhere. It organizes completed results and pending decisions into an email-style inbox. Earlier versions served more than 2,000 people inside Amazon, supporting meeting preparation, email drafting, Slack summaries, CRM logging, and research. The public application was rebuilt as an open source project. Deployable: Yes. Pizza Bot offers macOS, Windows, and Linux desktop builds, and browser and terminal clients connected to a local or standalone backend. Its code is licensed under Apache 2.0. An Inbox for Asynchronous Work Pizza Bot separates tasks into All, the thread history; Unread, completed work awaiting review; and Action, work paused for approval or an answer. Users can organize threads into folders and inspect delegated workers in the Activity panel. Tasks can start manually, through cron schedules, or through webhooks. The server owns scheduling. After downtime, missed cron intervals produce 1 catch-up run instead of replaying every missed interval. Trigger occurrences are recorded durably. How the Runtime Works The application uses DeepAgents and LangGraph for stateful execution. A Hono API server owns runtime execution and storage. Electron and browser clients share a React interface, while all clients communicate with the server over HTTP and server-sent events. LangGraph checkpoints retain thread state and approval pauses; separate SQLite stores hold cross-thread memory and application metadata. Reconnecting clients can replay buffered events. Closing a thread or disconnecting a client does not stop a running server. However, quitting the desktop app stops its embedded server and ends active runs. Checkpoints preserve the thread, but the step in flight can be lost. An always-on backend is required for work to continue after that desktop app exits. Skills, Tools, and Approval Controls Pizza Bot supports Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter, and Ollama. Configure a provider under Settings > Providers before running tasks. The agent has scratch-file operations and a sandboxed JavaScript interpreter without network or host-filesystem access. It can delegate through task when ready skill workers exist. The filesystem layer separately supports explicit folder grants and persistent memory. MCP servers expose external tools. Each SKILL.md defines a worker’s instructions and scoped tool access. A skill becomes callable only when its declared dependencies are available. Existing Claude Code-compatible .mcp.json configurations are supported, and plugins package skills with MCP servers. Skill authors configure interruptOn and allowedDecisions to require approval for specific tools. Depending on that policy, users can approve, edit proposed arguments, or reject an action. These controls must be configured for the relevant tools. Interactive Explainer Run the illustrative custom-skill workflow below. Compare an always-on backend with an embedded desktop server, close the client during execution, and approve, edit, or reject the proposed action. Animation timing is illustrative; no external actions occur. PIZZA BOT / BACKGROUND WORKIllustrative custom skill Backend locationAlways-on serverEmbedded in desktop Run example Close desktop Client: OpenServer: RunningCheckpoint: None WorkerApprovalResult Start a research brief, then close the desktop while it runs. All 0 Unread 0 Action 0 Prepare a research brief The custom skill gates its publish tool with interruptOn. publish_brief({title: “Research brief”}) Proposed title ApproveEditReject Read result Desktop closed Open the desktop to review its inbox. Simulation only. No external calls.Approval specification Key Takeaways An inbox for background agents: Pizza Bot organizes task history in All, completed work in Unread, and requests for approval or input in Action. Tasks support manual, cron, and webhook triggers. Persistent execution with DeepAgents and LangGraph: Checkpoints preserve thread state and approval pauses. Tasks continue after client disconnection while the backend remains running. Multiple model providers: Pizza Bot supports Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter, and local models through Ollama. Configured providers and tools can receive task data. Scoped skills and configurable approvals: MCP servers expose tools, while SKILL.md files define specialist workers. Tool-specific policies let users approve, edit, or reject proposed actions. Self-hosted and deployable: Apache 2.0 code, desktop builds, and standalone backend options are available. Each SQLite data directory supports 1 backend process. Check out the Technical details and GitHub Repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents appeared first on MarkTechPost.

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

An agent, in its simplest form, is an LLM calling tools in a loop. That loop works for short jobs. Give it a task that runs for an hour and 200 tool calls, and it breaks in 2 predictable ways. The AWS Samples design guide for autonomous cloud coding agents names them directly: shallow agents suffer from context overflow, get distracted (goal loss), and do not maintain state over long periods. The layer that fixes this is not the model. It is the harness, which AWS describes as managing everything but the model. This article opens up that layer. Compaction, memory strategy, context budgeting, and todo-state are the machinery that turns a shallow loop into a deep agent. We look at how LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore implement each one, with the actual thresholds they ship. Why a bigger window does not fix it The obvious fix is a larger context window. The evidence says it helps less than expected. Chroma’s Context Rot report evaluated 18 LLMs, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found that performance grows increasingly unreliable as input length grows, even on simple retrieval tasks. Anthropic’s context engineering guide explains the mechanism: attention creates n² pairwise relationships for n tokens, so every added token depletes a finite “attention budget.” Context is a resource with diminishing returns, not a bucket. For an agent loop, this is worse than it sounds. Manus reports that a typical task needs around 50 tool calls, and that the input-to-output token ratio runs near 100:1. Each observation lands in context and stays there. The original instruction drifts toward the middle of the window, which is exactly where recall degrades. Goal loss is not only a model bug. It is the expected outcome of an unmanaged context on a long enough task. Mechanism 1: Context budgeting and offloading The first job of a harness is deciding what never enters the window at all. Deep Agents ships 2 offloading rules with hard numbers. When a tool response exceeds 20,000 tokens, it is written to the filesystem and replaced with a file path plus a preview of the first 10 lines. When session context crosses 85% of the model’s window, older write and edit tool calls, whose full file contents already live on disk, are truncated to a pointer. Only after offloading runs out of room does the harness fall back to summarization. Claude Code applies the same budgeting to what loads before the first prompt. Auto memory is capped at the first 200 lines or 25KB. MCP tool schemas stay deferred by default, with only tool names listed, and full schemas load on demand via tool search. After compaction, any re-read file over 5,000 tokens comes back as a path reference rather than content. The context window simulation in the Claude Code docs makes the payoff concrete: a research subagent reads 6,100 tokens of files and returns a 420-token result to the parent. That subagent pattern is budgeting at the architecture level. Anthropic’s guide notes that each subagent may burn tens of thousands of tokens exploring, but returns a distilled summary, often 1,000 to 2,000 tokens. The AWS AgentCore walkthrough builds exactly this: a coordinator spawns 3 browser subagents in parallel, each in its own MicroVM, and an analyst subagent receives only their structured findings. AWS reports a 4 to 6 minute expected runtime, and notes that sequential processing would take up to 3x longer. Mechanism 2: Compaction When offloading is not enough, the harness summarizes. Compaction is the practice of taking a conversation nearing the window limit, summarizing it, and reinitiating a new context with the summary. It is also where goal loss most often happens, because a lossy summary can drop the one constraint that mattered. The implementations differ in what they promise to keep. Claude Code’s compaction prompt preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool outputs. Right after compaction it re-reads up to 5 of the files modified most recently, reloads the rules matching those files, and re-injects invoked skill bodies, capped at 5,000 tokens per skill and 25,000 total. The docs are explicit that detailed instructions from early in the conversation may be lost, which is why persistent rules belong in the project-root CLAUDE.md, which is re-injected from disk. Users can steer the pass with /compact focus on the auth bug fix or move the trigger point with /autocompact. Deep Agents made goal preservation a structural feature. Its summary is a structured document with dedicated fields for session intent, artifacts created, and next steps. The LangChain team added those fields after forced-summarization experiments showed the change improved performance. The full original transcript is also written to the filesystem, so a fact that was summarized away can be recovered by read_file later. Compaction has moved into the API layer too. OpenAI’s Responses API offers server-side compaction via context_management with a compact_threshold, plus a standalone /responses/compact endpoint that returns a compacted context window containing an opaque encrypted compaction item; OpenAI instructs developers to pass that returned window unchanged into the next call. OpenAI says Codex relies on this mechanism to sustain long-running coding tasks. The Claude Developer Platform exposes a compact_20260112 context-management edit with custom instructions and a pause_after_compaction option for inserting content before the model continues. When you write custom instructions there, they replace the default prompt entirely, so a compaction prompt is a real engineering artifact, not a setting. Mechanism 3: Todo-state and recitation Compaction protects the goal at the moment of summarization. Todo-state protects it on every turn in between. Manus described the trick plainly: its agent creates a todo.md and rewrites it step by step, checking items off. Rewriting the list recites the objectives into the end of the context, pushing the global plan into the model’s recent attention span and reducing “lost in the middle” drift. No architecture change is required. It is natural language used to bias the model’s own attention.

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction

In this tutorial, we build an end-to-end hierarchical Neural Radiance Field (NeRF) using JAX, Flax, Optax, and the volume-rendering primitives provided by jax3d. We first construct a synthetic multi-view dataset from an analytic scene containing volumetric geometry and view-dependent radiance, using sample_along_rays and volume_rendering to establish the forward rendering process. We then implement a NeRF with positional encoding, skip connections, separate coarse and fine networks, and view-direction conditioning, followed by hierarchical importance sampling through sample_piecewise_constant_pdf. We train the model with JAX JIT compilation, Adam optimization, exponential learning-rate decay, and gradient clipping, and finally evaluate novel-view synthesis using PSNR, depth and opacity visualization, sampling diagnostics, 360-degree rendering, and marching-cubes geometry extraction. Copy CodeCopiedUse a different Browser import os, sys, subprocess, importlib.util, functools, dataclasses, time, math def _sh(cmd): subprocess.run(cmd, shell=True, check=False, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) print(“Installing dependencies …”) _sh(f'{sys.executable} -m pip install -q “etils[array-types,epy,etree,enp]” ‘ f’chex flax optax scikit-image’) REPO_DIR = “/content/jax3d” if os.path.isdir(“/content”) else os.path.abspath(“./jax3d”) if not os.path.isdir(REPO_DIR): print(“Cloning google-research/jax3d …”) _sh(f”git clone -q –depth 1 https://github.com/google-research/jax3d.git {REPO_DIR}”) def _load_module_by_path(name, path): “””Load a single .py file without triggering the parent package __init__. `from jax3d.math import volume_rendering` also works if you run `pip install .` inside the clone, but that pulls in gin/tfds/etc. “”” spec = importlib.util.spec_from_file_location(name, path) mod = importlib.util.module_from_spec(spec) sys.modules[name] = mod spec.loader.exec_module(mod) return mod _VR_PATH = os.path.join(REPO_DIR, “jax3d”, “jax3d”, “math”, “volume_rendering.py”) if not os.path.exists(_VR_PATH): _VR_PATH = os.path.join(REPO_DIR, “jax3d”, “math”, “volume_rendering.py”) try: j3vr = _load_module_by_path(“j3d_volume_rendering”, _VR_PATH) except Exception as e: raise SystemExit( f”Could not load {_VR_PATH}: {e}n” “Try: pip install -U ‘etils[array-types,epy,etree,enp]==1.9.4’ and re-run.” ) import numpy as np import jax import jax.numpy as jnp import flax.linen as nn import optax from flax.training import train_state import matplotlib.pyplot as plt from PIL import Image print(“jax”, jax.__version__, “| device:”, jax.devices()[0].device_kind, f”({jax.devices()[0].platform})”) print(“jax3d volume_rendering API:”, [n for n in (“sample_along_rays”, “volume_rendering”, “sample_piecewise_constant_pdf”, “sample_1d”) if hasattr(j3vr, n)]) @dataclasses.dataclass class Config: H: int = 64; W: int = 64 n_train_views: int = 24; n_test_views: int = 3 cam_radius: float = 3.2; fov_deg: float = 40.0 near: float = 1.9; far: float = 4.7 gt_samples: int = 256 n_coarse: int = 64; n_fine: int = 64 deg_pos: int = 10; deg_dir: int = 4 width: int = 128; depth: int = 6; skip: int = 3 batch_rays: int = 2048; steps: int = 2500 lr_init: float = 5e-4; lr_final: float = 5e-6 chunk: int = 4096 grid_res: int = 96 cfg = Config() if jax.devices()[0].platform == “cpu”: print(“n!! No GPU detected — switching to a small CPU-friendly config.”) print(” (Runtime > Change runtime type > T4 GPU for the full version.)n”) cfg = dataclasses.replace(cfg, H=40, W=40, n_train_views=14, steps=400, gt_samples=128, n_coarse=32, n_fine=32, width=64, depth=4, skip=2, batch_rays=1024, chunk=1600, grid_res=64) def _normalize(v, axis=-1): return v / (np.linalg.norm(v, axis=axis, keepdims=True) + 1e-9) def look_at(eye, target=(0., 0., 0.), up=(0., 0., 1.)): “””OpenGL/NeRF convention camera-to-world: +x right, +y up, camera looks at -z.””” eye, target, up = map(lambda a: np.asarray(a, np.float32), (eye, target, up)) fwd = _normalize(target – eye) right = _normalize(np.cross(fwd, up)) trueup = np.cross(right, fwd) c2w = np.eye(4, dtype=np.float32) c2w[:3, :3] = np.stack([right, trueup, -fwd], axis=1) c2w[:3, 3] = eye return c2w def orbit_poses(n, radius, elev_lo=18., elev_hi=58., phase=0.0): “””Golden-angle azimuths + monotone elevations => well-spread views on a dome.””” i = np.arange(n, dtype=np.float64) + 0.5 az = 2 * np.pi * ((i * 0.6180339887) + phase) elev = np.arcsin(np.linspace(np.sin(np.deg2rad(elev_lo)), np.sin(np.deg2rad(elev_hi)), n)) eyes = np.stack([radius * np.cos(elev) * np.cos(az), radius * np.cos(elev) * np.sin(az), radius * np.sin(elev)], axis=-1).astype(np.float32) return np.stack([look_at(e) for e in eyes], axis=0) def rays_from_pose(c2w, H, W, focal): “””Returns (origins, dirs) of shape [H, W, 3]; dirs are unit-length, so the depths returned by jax3d’s sampler are true world-space distances.””” i, j = np.meshgrid(np.arange(W, dtype=np.float32), np.arange(H, dtype=np.float32), indexing=”xy”) cam_dirs = np.stack([(i – W * .5 + .5) / focal, -(j – H * .5 + .5) / focal, -np.ones_like(i)], axis=-1) dirs = _normalize(cam_dirs @ c2w[:3, :3].T) origins = np.broadcast_to(c2w[:3, 3], dirs.shape) return origins.astype(np.float32).copy(), dirs.astype(np.float32) FOCAL = 0.5 * cfg.W / math.tan(0.5 * math.radians(cfg.fov_deg)) We set up the JAX3D environment, install the required dependencies, and load the volume_rendering module directly from the cloned repository. We configure GPU/CPU-adaptive training parameters and establish the camera model using pinhole intrinsics, look-at poses, and orbit-based camera placement. We then generate normalized world-space rays from each camera pose, providing the geometric foundation for the rendering pipeline. Copy CodeCopiedUse a different Browser LIGHT = jnp.asarray(_normalize(np.array([0.55, 0.75, 0.85], np.float32))) _SPHERES = [ (jnp.array([0.34, 0.02, -0.22]), 0.36, jnp.array([0.90, 0.24, 0.22])), (jnp.array([-0.32, 0.28, 0.05]), 0.26, jnp.array([0.25, 0.78, 0.36])), (jnp.array([-0.05, -0.36, 0.24]), 0.22, jnp.array([0.28, 0.40, 0.95])), ] def _sphere_field(pos, vdir, center, radius, albedo): d = pos – center dist = jnp.linalg.norm(d, axis=-1) n = d / (dist[…, None] + 1e-8) sigma = 80.0 * jax.nn.sigmoid((radius – dist) / 0.015) v = -vdir refl = 2.0 * jnp.sum(n * v, -1, keepdims=True) * n – v spec = 0.65 * jnp.clip(jnp.sum(refl * LIGHT, -1), 0., 1.) ** 24 lamb = 0.35 + 0.65 * jnp.clip(jnp.sum(n * LIGHT, -1), 0., 1.) rgb = jnp.clip(albedo * lamb[…, None] + spec[…, None], 0., 1.) return sigma, rgb def _floor_field(pos): x, y, z = pos[…, 0], pos[…, 1], pos[…, 2] m = (jax.nn.sigmoid((0.06 – jnp.abs(z + 0.62)) / 0.008) * jax.nn.sigmoid((0.85 – jnp.abs(x)) / 0.01) * jax.nn.sigmoid((0.85 – jnp.abs(y)) / 0.01)) checker = (jnp.floor(x * 3.0) + jnp.floor(y * 3.0)) % 2.0 rgb = jnp.where(checker[…, None] > 0.5, jnp.array([0.86, 0.86, 0.89]), jnp.array([0.22, 0.25, 0.30])) return 80.0 * m, rgb def gt_field(pos, vdir): “””pos, vdir: […, 3] -> (sigma […], rgb […, 3]). Density-weighted blend.””” sig_sum = 0.0 col_sum = 0.0 for c, r, a in _SPHERES: s, rgb = _sphere_field(pos, vdir, c, r, a) sig_sum = sig_sum + s col_sum = col_sum + s[…, None] * rgb s, rgb = _floor_field(pos) sig_sum = sig_sum + s col_sum = col_sum + s[…, None] * rgb return sig_sum, col_sum / (sig_sum[…, None] + 1e-8) WHITE_BG = jnp.ones((3,), jnp.float32) @jax.jit def render_ground_truth(origins, dirs): “””Fine-grained volumetric render of the analytic scene -> RGB + depth.””” depths, positions = j3vr.sample_along_rays( ray_origins=origins, ray_directions=dirs, near=cfg.near, far=cfg.far, sample_count=cfg.gt_samples, deterministic=True) vdir = jnp.broadcast_to(dirs[…, None, :],

Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

In most decoder-only LLMs, nothing computed at the last layer of token t feeds the first layer of token t+1; positions communicate only through attention over cached keys and values. A Princeton researcher’s (Yifan Zhang) technical report, Recurrent Looped Transformer (RLT), proposes closing that loop. The decoder’s final hidden state and its layerwise sliding-window attention (SWA) cache are carried into the next token, across both prompt and response, with no reset at the boundary. The proposed research is a design specification. It defines the architecture, execution schedules, and RL replay contract, and it explicitly reports no measured efficiency, reasoning quality, or scaling results. How RLT is Built Recurrent Looped Transformer (RLT) pairs a causal encoder with a recurrent decoder. The encoder processes tokens in parallel under a causal mask and produces representations e_t, from which key-value memory M≤t is projected; memory groups can be shared across decoder layers (G = 1) or kept layer-specific (G = L_D). The decoder holds the recurrence. Its complete state is Ht = (st, CtD), where st is the final decoder output and CtD holds the retained SWA keys and values at every decoder layer. For each token, a gated merge combines et with the previous output s{t-1}, then each decoder block runs causal SWA over decoder activations, cross-attention to encoder memory, and an FFN. The window W includes the current token, so at most W – 1 historical entries per layer are retained. The next-token distribution is read from st, and initialization happens once before BOS with a learned start state s* and an empty cache. The reference tied configuration uses 48 encoder and 48 decoder layers with compatible attention and FFN weights shared between them. Each token therefore executes 96 logical blocks, though decoder blocks add cross-attention, so per-block FLOPs are not equal. Zhang calls this parameter reuse, not activation copying. The 3 Design Principles Latent reasoning with unbounded temporal depth: After t processed tokens, the state path from s0 traverses t·LD decoder blocks, or 48t in the reference configuration. Per-token work stays fixed while the path’s structural depth grows with the sequence. The research report warns that gates and contraction may suppress long paths; structural depth is not a reasoning guarantee. Model-hardware co-design: Encoder features and memory projections for known tokens use token-parallel kernels. Decoder transitions stay sequential within a sequence, but ready updates from independent sequences can share one batched kernel. The report states plainly that no exact parallel scan is assumed for the nonlinear decoder, no reduced-prefill speedup is claimed, and a standard parallel SWA decoder pass is not equivalent to the recurrence. Batching, kernel fusion, and checkpointing are listed as implementation targets, not completed kernels. Model-RL algorithm co-design: Pretraining, SFT, sampling, and RL replay share one state transition. For RL, the sampler records each action’s behavior log-probability under its actual sampling distribution, including temperature and truncation. The trainer rebuilds encoder memory, the recurrent output, and every SWA cache from the sequence start under current parameters before scoring each action; old rollout states are never reused. Proposition 3.1 formalizes the payoff: moving the prompt-response split leaves the conditional distribution unchanged for a fixed token history. Training and Serving Pretraining is full-sequence next-token prediction with full backpropagation through time. SFT masks the loss to assistant targets but never masks state updates, so assistant losses backpropagate through user and tool tokens. Appendix B shows why partial detaching is risky: the state-to-state Jacobian has cross terms through decoder KV, so detaching only st leaves gradient paths through the cache; any truncated-BPTT scheme must name every detached tensor. For multi-turn serving, an exact prefix snapshot includes encoder cache and memory, the complete decoder state, position metadata, the window convention, and model version. A fixed-weight snapshot can be reused because the state is independent of the serving split; weight updates invalidate old states, and editing a prefix forces recomputation from an earlier checkpoint. External tokens in multi-turn RL update the state but get no importance-ratio factors. How It Relates to Prior Work Encoder-derived memory follows YOCO, which caches KV once for a cross-decoder, and DeepSeek-V4.1-Flash, which projects decoder global KV from final encoder states; RLT keeps the memory but drops prompt-wide decoder skipping. Temporal feedback builds on Feedback Transformer and Recurrent Transformer; RLT instead feeds the previous final decoder output into the next decoder input and runs recurrence over the prompt too. Depth-wise reuse connects to Universal Transformers and recurrent-depth latent reasoning; the replay argument extends Zhang’s prefill-decode kernel mismatch note. Interactive Explainer Key Takeaways RLT carries the full decoder state (final output plus layerwise SWA cache) across every prompt and response token with no boundary reset. Reference config: 48 tied encoder and decoder layers, 96 logical blocks per token, state path of 48t blocks after t tokens. Hardware opportunities: encoder parallelism and batching across sequences; no parallel scan or reduced-prefill speedup is claimed. RL replay rebuilds all states under current parameters while keeping recorded behavior log-probabilities as ratio denominators. No measured results: reasoning quality, efficiency, and RL scaling remain open validation targets. Check out the Technical Report, GitHub repository, and Project Page. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth appeared first on MarkTechPost.

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth Leggi l'articolo »

We use cookies to improve your experience and performance on our website. You can learn more at Politica sulla privacy and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
it_IT