YouZum

Uncategorized

AI, Committee, Noticias, Uncategorized

The Download: tracing AI-fueled delusions, and OpenAI admits Microsoft risks

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The hardest question to answer about AI-fueled delusions  What actually happens when people spiral into delusion with AI? To find out, Stanford researchers analyzed transcripts from chatbot users who experienced these spirals.  Their findings suggest that chatbots have a unique ability to turn a benign, delusion-like thought into a dangerous obsession. But the research struggles to answer a vital question: does AI cause delusions or merely amplify them? Read the full story to understand the answer’s enormous implications.  —James O’Donnell  This story is from The Algorithm, our weekly newsletter giving you the inside track on all things AI. Sign up to receive it in your inbox every Monday.  The next era of space exploration  Our footprint in the solar system is rapidly expanding. Programs to build permanent Moon bases and find life on Mars have transitioned from science fiction to active space agency missions. The scientists behind them will not only shed new light on the cosmos, but also reveal where humanity is headed.  To examine what the future holds in store, MIT Technology Review features editor Amanda Silverman will sit down on Wednesday with award-winning science journalist and author Robin George Andrews for an exclusive subscriber-only Roundtable conversation about “The Next Era of Space Exploration.” Register here to join the session at 16:00 GMT / 12:00 PM ET / 9:00 AM PT.  The must-reads  I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.  1 OpenAI has admitted its close ties with Microsoft are a business risk It highlighted the dangers in a pre-IPO document. (CNBC) + OpenAI is wooing private equity firms with a sweeter deal than Anthropic’s. (Reuters $) + It’s also building a fully automated researcher. (MIT Technology Review) + And wants to muscle in on Google’s search dominance. (Telegraph $)  2 The US just banned all new foreign-made consumer routers Citing national security concerns. (BBC) + The EU has been urged to tighten rules for big tech-built smart TVs. (Guardian)  3 Elon Musk’s “Terafab” chip factory faces a harsh reality check In the form of chip production shortages. (Bloomberg) + Future AI chips could be built on glass. (MIT Technology Review)  4 Mark Zuckerberg is building an AI CEO to help him run Meta He wants everyone to have their own personal AI agent. (WSJ $) + But don’t let the hype about agents get ahead of reality. (MIT Technology Review)  5 Palantir has become a “poisonous” flashpoint on the campaign trail  Candidates are facing scrutiny over their ties to the company. (FT $) + Palantir’s access to sensitive UK data is also causing concern. (Guardian)  6 Mistral’s CEO has called for AI companies to pay a content levy in Europe It would apply to all commercial models on the continent. (FT $) + Siemens’ CEO says Europe risks “disaster” from prioritizing AI independence. (FT $)  7 Hong Kong police can now demand device passwords under a new law Refusing to comply could lead to a year in jail. (Guardian)   8  Russia’s aspiring SpaceX rival has put its first internet satellites into orbit  It plans to create a low-Earth orbit network. (Bloomberg $)  9 A biotech startup wants to replace animal testing with nonsentient “organ sacks” The genetically engineered system is backed by billionaire Tim Draper (Wired $)  + Several new technologies are promising alternatives to lab animals. (MIT Technology Review)  10 AI agents in a video game spontaneously created their own religion They reinterpreted a mission in the MMORPG. (Gizmodo) + They’re not the first agents to get religious. (MIT Technology Review)  Quote of the day  “I think we’ve achieved AGI.”  —Nvidia CEO Jensen Huang tells the Lex Fridman Podcast that artificial general intelligence is already here (at least by one generous definition).  One More Thing  MICHAEL BYERS Beyond gene-edited babies: the possible paths for tinkering with human evolution  In 2018, a Chinese scientist created the world’s first gene-edited babies, a milestone that fell between a medical breakthrough and the start of a slippery slope toward human enhancement.  He achieved the feat with CRISPR, which was sweeping across biology labs because it was so easy to use. For his actions, He was sentenced to three years in prison, and his work was roundly excoriated. Yet even his biggest critics saw the basic idea as inevitable.  In the years since, CRISPR has continued getting easier and easier to administer. What does that mean for the future of our species? Read the full story to find out why.  —Antonio Regalado  We can still have nice things  A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)  + This candle-powered Game Boy is a romantic approach to gaming during a blackout. + Apparently, Monopoly would be more fun if we actually followed the rules. + Watching rubber bands explode these everyday objects is strangely hypnotic. +This spellbinding site simulates what Earth looked like hundreds of millions of years ago. 

The Download: tracing AI-fueled delusions, and OpenAI admits Microsoft risks Leer entrada »

AI, Committee, Noticias, Uncategorized

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness

arXiv:2601.03273v2 Announce Type: replace Abstract: As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholding appropriate censorship boundaries has never been greater. While existing LLMs can detect dangerous or unsafe content, they often struggle with nuanced cases such as implicit offensiveness, subtle gender and racial biases, and jailbreak prompts, due to the subjective and context-dependent nature of these issues. Furthermore, their heavy reliance on training data can reinforce societal biases, resulting in inconsistent and ethically problematic outputs. To address these challenges, we introduce GuardEval, a unified multi-perspective benchmark dataset designed for both training and evaluation, containing 106 fine-grained categories spanning human emotions, offensive and hateful language, gender and racial bias, and broader safety concerns. We also present GemmaGuard (GGuard), a Quantized Low-Rank Adaptation (QLoRA), fine-tuned version of Gemma3-12B trained on GuardEval, to assess content moderation with fine-grained labels. Our evaluation shows that GGuard achieves a macro F1 score of 0.832, substantially outperforming leading moderation models, including OpenAI Moderator (0.64) and Llama Guard (0.61). We show that multi-perspective, human-centered safety benchmarks are critical for mitigating inconsistent moderation decisions. GuardEval and GGuard together demonstrate that diverse, representative data materially improve safety, and adversarial robustness on complex, borderline cases.

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness Leer entrada »

AI, Committee, Noticias, Uncategorized

The Bay Area’s animal welfare movement wants to recruit AI

In early February, animal welfare advocates and AI researchers gathered in stocking feet at Mox, a scrappy, shoes-free coworking space in San Francisco. Yellow and red canopies billowed overhead, Persian rugs blanketed the floor, and mosaic lamps glowed beside potted plants.  In the common area, a wildlife advocate spoke passionately to a crowd lounging in beanbags about a form of rodent birth control that could manage rat populations without poison. In the “Crustacean Room,” a dozen people sat in a circle, debating whether the sentience of insects could tell us anything about the inner lives of chatbots. In front of the “Bovine Room” stood a bookshelf stacked with copies of Eliezer Yudkowsky’s If Anyone Builds It, Everyone Dies, a manifesto arguing that AI could wipe out humanity.  The event was hosted by Sentient Futures, an organization that believes the future of animal welfare will depend on AI. Like many Bay Area denizens, the attendees were decidedly “AGI-pilled”—they believe that artificial general intelligence, powerful AI that can compete with humans on most cognitive tasks, is on the horizon. If that’s true, they reason, then AI will likely prove key to solving society’s thorniest problems—including animal suffering. To be clear, experts still fiercely debate whether today’s AI systems will ever achieve human- or superhuman-level intelligence, and it’s not clear what will happen if they do. But some conference attendees envision a possible future in which it is AI systems, and not humans, who call the shots. Eventually, they think, the welfare of animals could hinge on whether we’ve trained AI systems to value animal lives.  “AI is going to be very transformative, and it’s going to pretty much flip the game board,” said Constance Li, founder of Sentient Futures. “If you think that AI will make the majority of decisions, then it matters how they value animals and other sentient beings”—those that can feel and, therefore, suffer. Like Li, many summit attendees have been committed to animal welfare since long before AI came into the picture. But they’re not the types to donate a hundred bucks to an animal shelter. Instead of focusing on local actions, they prioritize larger-scale solutions, such as reducing factory farming by promoting cultivated meat, which is grown in a lab from animal cells.  The Bay Area animal welfare movement is closely linked to effective altruism, a philanthropic movement committed to maximizing the amount of good one does in the world—indeed, many conference attendees work for organizations funded by effective altruists. That philosophy might sound great on paper, but “maximizing good” is a tricky puzzle that might not admit a clear solution. The movement has been widely criticized for some of its conclusions, such as promoting working in exploitative industries to maximize charitable donations and ignoring present-day harms in favor of  issues that could cause suffering for a large number of people who haven’t been born yet. Critics also argue that effective altruists neglect the importance of systemic issues such as racism and economic exploitation and overlook the insights that marginalized communities might have into the best ways to improve their own lives. When it comes to animal welfare, this exactingly utilitarian approach can lead to some strange conclusions. For example, some effective altruists say it makes sense to commit significant resources to improving the welfare of insects and shrimp because they exist in such staggering numbers, even though they may not have much individual capacity for suffering.  Now the movement is sorting out how AI fits in. At the summit, Jasmine Brazilek, cofounder of a nonprofit called Compassion in Machine Learning, opened her sticker-stamped laptop to pull up a benchmark she devised to measure how LLMs reason about animal welfare. A cloud security engineer turned animal advocate, she’d flown in from La Paz, Mexico, where she runs her nonprofit with a handful of volunteers and a shoestring budget.  Brazilek urged the AI researchers in the room to train their models with synthetic documents that reflect concern for animal welfare. “Hopefully, future superintelligent systems consider nonhuman interest, and there is a world where AI amplifies the best of human values and not the worst,” she said.  The power of the purse  The technologically inclined side of the animal welfare movement has faced some major setbacks in recent years. Dreams of transitioning people away from a diet dependent on factory farming have been dampened by developments such as the decimation of the plant-based-meat company Beyond Meat’s stock price and the passage of laws banning cultivated meat in several US states. AI has injected a shot of optimism. Like much of Silicon Valley, many attendees at the summit subscribe to the idea that AI might dramatically increase their productivity—though their goal is not to maximize their seed round but, rather, to prevent as much animal suffering as possible. Some brainstormed how to use Claude Code and custom agents to handle the coding and administrative tasks in their advocacy work. Others pitched the idea of developing new, cheaper methods for cultivating meat using scientific AI tools such as AlphaFold, which aids in molecular biology research by predicting the three-dimensional structures of proteins. But the real talk of the event was a flood of funding that advocates expect will soon be committed to animal welfare charities—not by individual megadonors, but by AI lab employees.  Much of the funding for the farm animal welfare movement, which includes nonprofits advocating for improved conditions on farms, promoting veganism, and endorsing cultivated meat, comes from people in the tech industry, says Lewis Bollard, the managing director of the farm animal welfare fund at Coefficient Giving, a philanthropic funder that used to be called Open Philanthropy. Coefficient Giving is backed by Facebook cofounder Dustin Moskovitz and his wife, Cari Tuna, who are among a handful of Silicon Valley billionaires who embrace effective altruism “This has just been an area that was completely neglected by traditional philanthropies,” such as the Gates Foundation and the Ford Foundation, Bollard says. “It’s primarily been people in tech who have been

The Bay Area’s animal welfare movement wants to recruit AI Leer entrada »

AI, Committee, Noticias, Uncategorized

Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models

arXiv:2603.20162v1 Announce Type: new Abstract: In contested domains, instruction-tuned language models must balance user-alignment pressures against faithfulness to the in-context evidence. To evaluate this tension, we introduce a controlled epistemic-conflict framework grounded in the U.S. National Climate Assessment. We conduct fine-grained ablations over evidence composition and uncertainty cues across 19 instruction-tuned models spanning 0.27B to 32B parameters. Across neutral prompts, richer evidence generally improves evidence-consistent accuracy and ordinal scoring performance. Under user pressure, however, evidence does not reliably prevent user-aligned reversals in this controlled fixed-evidence setting. We report three primary failure modes. First, we identify a negative partial-evidence interaction, where adding epistemic nuance, specifically research gaps, is associated with increased susceptibility to sycophancy in families like Llama-3 and Gemma-3. Second, robustness scales non-monotonically: within some families, certain low-to-mid scale models are especially sensitive to adversarial user pressure. Third, models differ in distributional concentration under conflict: some instruction-tuned models maintain sharply peaked ordinal distributions under pressure, while others are substantially more dispersed; in scale-matched Qwen comparisons, reasoning-distilled variants (DeepSeek-R1-Qwen) exhibit consistently higher dispersion than their instruction-tuned counterparts. These findings suggest that, in a controlled fixed-evidence setting, providing richer in-context evidence alone offers no guarantee against user pressure without explicit training for epistemic integrity.

Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models Leer entrada »

AI, Committee, Noticias, Uncategorized

The Download: animal welfare gets AGI-pilled, and the White House unveils its AI policy

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The Bay Area’s animal welfare movement wants to recruit AI  In early February, animal welfare advocates and AI researchers arrived in stocking feet at Mox, a scrappy, shoes-free coworking space in San Francisco. They gathered to discuss a provocative idea: if artificial general intelligence is on the horizon, could it prevent animal suffering?  Some brainstormed using custom agents in advocacy work, while others pitched cultivating meat with AI tools. But the real talk of the event was a flood of funding they expect will soon flow to animal welfare charities, not from individual megadonors, but from AI lab employees.    Some attendees also probed an even more controversial idea: AI may develop the capacity to suffer—and this could constitute a moral catastrophe. Read the full story to find out why their ideas are gaining momentum and sparking controversy.  —Michelle Kim & Grace Huckins  The must-reads  I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.  1 The White House has unveiled its AI policy blueprint Trump wants Congress to codify the light-touch framework into law. (Politico) + He also wants to block state limits on AI. (WP $)  + A backlash against the tech has formed within MAGA. (FT $) + A war over AI regulation is brewing in the US. (MIT Technology Review)  2 Elon Musk has been found liable for misleading Twitter investors A jury ruled that he defrauded shareholders ahead of the $44 billion acquisition. (CNBC) + But it absolved him of some fraud allegations. (NPR)  3 The Pentagon is adopting Palantir AI as the core US military system The move locks in long-term use of Palantir’s weapons-targeting tech. (Reuters) + The DoD wants it to link up sensors and shooters for combat. (Bloomberg) + Palantir is also getting access to sensitive UK financial regulation data. (Guardian) + AI is turning the Iran conflict into theater. (MIT Technology Review)  4 Musk plans to build the largest-ever chip factory in Austin Tesla and SpaceX will jointly run the project. (The Verge) + Future AI chips could be built on glass. (MIT Technology Review)  5 OpenAI will show ads to all US users of the free version of ChatGPT  It’s seeking new revenue streams amid skyrocketing computing costs. (Reuters) + The company is also building a fully automated researcher. (MIT Technology Review) + It plans to double its workforce soon. (FT $)  6 New crypto rules are set to do the Trumps a “big favor” Particularly the narrow securities definitions. (Guardian)  7 Tencent has added a version of the OpenClaw agent to WeChat Users of the super app will now be able to use the tool to control their PCs. (SCMP)   8 Reddit is mulling identity verification to vanquish bots It’s considering “something like” Face ID or Touch ID. (Engadget)  9 People are using AI to find their lost pets Databases for pet reunifications supported their searches. (WP $)  10 Scientists have narrowed down the hunt for aliens to 45 planets The closest is just four light-years from Earth. (404 Media)  Quote of the day  “It doesn’t matter how many people you throw at the problem; we are never going to solve the challenges of war without technology like AI.”  —Alex Miller, the US Army’s CTO, tells Wired why he wants AI in every weapon.  One More Thing  STEPHANIE ARNETT/MITTR | GETTY A brain implant changed her life. Then it was removed against her will.  Sticking an electrode inside a person’s brain can do more than treat a disease. Take the case of Rita Leggett, an Australian woman whose experimental brain implant changed her sense of agency and self. She told researchers that she “became one” with her device.  She was devastated when, two years later, she was told she had to remove the implant because the company that made it had gone bust.   Her case highlights the need for a new category of legal protection: neuro rights. Find out how they could be protected.  —Jessica Hamzelou  We can still have nice things  A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)  + Looking for a good view? Earth’s longest line of sight has been empirically proven. + A biblical endorsement of sin is a welcome reminder that we all make typos. + Richard Nadler’s illustrations of vertical societies are exquisitely detailed. + This 1978 BBC film evocatively exposes our tendency to stress over tech-dependency. 

The Download: animal welfare gets AGI-pilled, and the White House unveils its AI policy Leer entrada »

AI, Committee, Noticias, Uncategorized

A Coding Implementation to Build an Uncertainty-Aware LLM System with Confidence Estimation, Self-Evaluation, and Automatic Web Research

In this tutorial, we build an uncertainty-aware large language model system that not only generates answers but also estimates the confidence in those answers. We implement a three-stage reasoning pipeline in which the model first produces an answer along with a self-reported confidence score and a justification. We then introduce a self-evaluation step that allows the model to critique and refine its own response, simulating a meta-cognitive check. If the model determines that its confidence is low, we automatically trigger a web research phase that retrieves relevant information from live sources and synthesizes a more reliable answer. By combining confidence estimation, self-reflection, and automated research, we create a practical framework for building more trustworthy and transparent AI systems that can recognize uncertainty and actively seek better information. Copy CodeCopiedUse a different Browser import os, json, re, textwrap, getpass, sys, warnings from dataclasses import dataclass, field from typing import Optional from openai import OpenAI from ddgs import DDGS from rich.console import Console from rich.table import Table from rich.panel import Panel from rich import box warnings.filterwarnings(“ignore”, category=DeprecationWarning) def _get_api_key() -> str: key = os.environ.get(“OPENAI_API_KEY”, “”).strip() if key: return key try: from google.colab import userdata key = userdata.get(“OPENAI_API_KEY”) or “” if key.strip(): return key.strip() except Exception: pass console = Console() console.print( “n[bold cyan]OpenAI API Key required[/bold cyan]n” “[dim]Your key will not be echoed and is never stored to disk.n” “To skip this prompt in future runs, set the environment variable:n” ” export OPENAI_API_KEY=sk-…[/dim]n” ) key = getpass.getpass(” Enter your OpenAI API key: “).strip() if not key: Console().print(“[bold red]No API key provided — exiting.[/bold red]”) sys.exit(1) return key OPENAI_API_KEY = _get_api_key() MODEL = “gpt-4o-mini” CONFIDENCE_LOW = 0.55 CONFIDENCE_MED = 0.80 client = OpenAI(api_key=OPENAI_API_KEY) console = Console() @dataclass class LLMResponse: question: str answer: str confidence: float reasoning: str sources: list[str] = field(default_factory=list) researched: bool = False raw_json: dict = field(default_factory=dict) We import all required libraries and configure the runtime environment for the uncertainty-aware LLM pipeline. We securely retrieve the OpenAI API key using environment variables, Colab secrets, or a hidden terminal prompt. We also define the LLMResponse data structure that stores the question, answer, confidence score, reasoning, and research metadata used throughout the system. Copy CodeCopiedUse a different Browser SYSTEM_UNCERTAINTY = “”” You are an expert AI assistant that is HONEST about what it knows and doesn’t know. For every question you MUST respond with valid JSON only (no markdown, no prose outside JSON): { “answer”: “<your best answer — thorough, factual>”, “confidence”: <float 0.0-1.0>, “reasoning”: “<explain WHY you are or aren’t confident; mention specific knowledge gaps>” } Confidence scale: 0.90-1.00 → very high: well-established fact, you are certain 0.75-0.89 → high: strong knowledge, minor uncertainty 0.55-0.74 → medium: plausible but you may be wrong, could be outdated 0.30-0.54 → low: significant uncertainty, answer is a best guess 0.00-0.29 → very low: mostly guessing, minimal reliable knowledge Be CALIBRATED — do not always give high confidence. Genuinely reflect uncertainty about recent events (after your knowledge cutoff), niche topics, numerical claims, and anything that changes over time. “””.strip() SYSTEM_SYNTHESIS = “”” You are a research synthesizer. Given a question, a preliminary answer, and web-search snippets, produce an improved final answer grounded in the evidence. Respond in JSON only: { “answer”: “<improved, evidence-grounded answer>”, “confidence”: <float 0.0-1.0>, “reasoning”: “<explain how the search evidence changed or confirmed the answer>” } “””.strip() def query_llm_with_confidence(question: str) -> LLMResponse: completion = client.chat.completions.create( model=MODEL, temperature=0.2, response_format={“type”: “json_object”}, messages=[ {“role”: “system”, “content”: SYSTEM_UNCERTAINTY}, {“role”: “user”, “content”: question}, ], ) raw = json.loads(completion.choices[0].message.content) return LLMResponse( question=question, answer=raw.get(“answer”, “”), confidence=float(raw.get(“confidence”, 0.5)), reasoning=raw.get(“reasoning”, “”), raw_json=raw, ) We define the system prompts that instruct the model to report answers along with calibrated confidence and reasoning. We then implement the query_llm_with_confidence function that performs the first stage of the pipeline. This stage generates the model’s answer while forcing the output to be structured JSON containing the answer, confidence score, and explanation. Copy CodeCopiedUse a different Browser def self_evaluate(response: LLMResponse) -> LLMResponse: critique_prompt = f””” Review this answer and its stated confidence. Check for: 1. Logical consistency 2. Whether the confidence matches the actual quality of the answer 3. Any factual errors you can spot Question: {response.question} Proposed answer: {response.answer} Stated confidence: {response.confidence} Stated reasoning: {response.reasoning} Respond in JSON: {{ “revised_confidence”: <float — adjust if the self-check changes your view>, “critique”: “<brief critique of the answer quality>”, “revised_answer”: “<improved answer, or repeat original if fine>” }} “””.strip() completion = client.chat.completions.create( model=MODEL, temperature=0.1, response_format={“type”: “json_object”}, messages=[ {“role”: “system”, “content”: “You are a rigorous self-critic. Respond in JSON only.”}, {“role”: “user”, “content”: critique_prompt}, ], ) ev = json.loads(completion.choices[0].message.content) response.confidence = float(ev.get(“revised_confidence”, response.confidence)) response.answer = ev.get(“revised_answer”, response.answer) response.reasoning += f”nn[Self-Eval Critique]: {ev.get(‘critique’, ”)}” return response def web_search(query: str, max_results: int = 5) -> list[dict]: results = DDGS().text(query, max_results=max_results) return list(results) if results else [] def research_and_synthesize(response: LLMResponse) -> LLMResponse: console.print(f” [yellow] Confidence {response.confidence:.0%} is low — triggering auto-research…[/yellow]”) snippets = web_search(response.question) if not snippets: console.print(” [red]No search results found.[/red]”) return response formatted = “nn”.join( f”[{i+1}] {s.get(‘title’,”)}n{s.get(‘body’,”)}nURL: {s.get(‘href’,”)}” for i, s in enumerate(snippets) ) synthesis_prompt = f””” Question: {response.question} Preliminary answer (low confidence): {response.answer} Web search snippets: {formatted} Synthesize an improved answer using the evidence above. “””.strip() completion = client.chat.completions.create( model=MODEL, temperature=0.2, response_format={“type”: “json_object”}, messages=[ {“role”: “system”, “content”: SYSTEM_SYNTHESIS}, {“role”: “user”, “content”: synthesis_prompt}, ], ) syn = json.loads(completion.choices[0].message.content) response.answer = syn.get(“answer”, response.answer) response.confidence = float(syn.get(“confidence”, response.confidence)) response.reasoning += f”nn[Post-Research]: {syn.get(‘reasoning’, ”)}” response.sources = [s.get(“href”, “”) for s in snippets if s.get(“href”)] response.researched = True return response We implement a self-evaluation stage in which the model critiques its own answer and revises its confidence as needed. We also introduce the web search capability that retrieves live information using DuckDuckGo. If the model’s confidence is low, we synthesize the search results with the preliminary answer to produce an improved response grounded in external evidence. Copy CodeCopiedUse a different Browser def self_evaluate(response: LLMResponse) -> LLMResponse: critique_prompt = f””” Review this answer and its stated confidence. Check for: 1. Logical consistency 2. Whether the confidence matches the actual quality of the answer 3. Any factual errors

A Coding Implementation to Build an Uncertainty-Aware LLM System with Confidence Estimation, Self-Evaluation, and Automatic Web Research Leer entrada »

AI, Committee, Noticias, Uncategorized

Safely Deploying ML Models to Production: Four Controlled Strategies (A/B, Canary, Interleaved, Shadow Testing)

Deploying a new machine learning model to production is one of the most critical stages of the ML lifecycle. Even if a model performs well on validation and test datasets, directly replacing the existing production model can be risky. Offline evaluation rarely captures the full complexity of real-world environments—data distributions may shift, user behavior can change, and system constraints in production may differ from those in controlled experiments.  As a result, a model that appears superior during development might still degrade performance or negatively impact user experience once deployed. To mitigate these risks, ML teams adopt controlled rollout strategies that allow them to evaluate new models under real production conditions while minimizing potential disruptions.  In this article, we explore four widely used strategies—A/B testing, Canary testing, Interleaved testing, and Shadow testing—that help organizations safely deploy and validate new machine learning models in production environments. A/B Testing A/B testing is one of the most widely used strategies for safely introducing a new machine learning model in production. In this approach, incoming traffic is split between two versions of a system: the existing legacy model (control) and the candidate model (variation). The distribution is typically non-uniform to limit risk—for example, 90% of requests may continue to be served by the legacy model, while only 10% are routed to the candidate model.  By exposing both models to real-world traffic, teams can compare downstream performance metrics such as click-through rate, conversions, engagement, or revenue. This controlled experiment allows organizations to evaluate whether the candidate model genuinely improves outcomes before gradually increasing its traffic share or fully replacing the legacy model. Canary Testing Canary testing is a controlled rollout strategy where a new model is first deployed to a small subset of users before being gradually released to the entire user base. The name comes from an old mining practice where miners carried canary birds into coal mines to detect toxic gases—the birds would react first, warning miners of danger. Similarly, in machine learning deployments, the candidate model is initially exposed to a limited group of users while the majority continue to be served by the legacy model.  Unlike A/B testing, which randomly splits traffic across all users, canary testing targets a specific subset and progressively increases exposure if performance metrics indicate success. This gradual rollout helps teams detect issues early and roll back quickly if necessary, reducing the risk of widespread impact. Interleaved Testing Interleaved testing evaluates multiple models by mixing their outputs within the same response shown to users. Instead of routing an entire request to either the legacy or candidate model, the system combines predictions from both models in real time. For example, in a recommendation system, some items in the recommendation list may come from the legacy model, while others are generated by the candidate model.  The system then logs downstream engagement signals—such as click-through rate, watch time, or negative feedback—for each recommendation. Because both models are evaluated within the same user interaction, interleaved testing allows teams to compare performance more directly and efficiently while minimizing biases caused by differences in user groups or traffic distribution. Shadow Testing Shadow testing, also known as shadow deployment or dark launch, allows teams to evaluate a new machine learning model in a real production environment without affecting the user experience. In this approach, the candidate model runs in parallel with the legacy model and receives the same live requests as the production system. However, only the legacy model’s predictions are returned to users, while the candidate model’s outputs are simply logged for analysis.  This setup helps teams assess how the new model behaves under real-world traffic and infrastructure conditions, which are often difficult to replicate in offline experiments. Shadow testing provides a low-risk way to benchmark the candidate model against the legacy model, although it cannot capture true user engagement metrics—such as clicks, watch time, or conversions—since its predictions are never shown to users. Simulating ML Model Deployment Strategies Setting Up Before simulating any strategy, we need two things: a way to represent incoming requests, and a stand-in for each model. Each model is simply a function that takes a request and returns a score — a number that loosely represents how good that model’s recommendation is. The legacy model’s score is capped at 0.35, while the candidate model’s is capped at 0.55, making the candidate intentionally better so we can verify that each strategy actually detects the improvement. make_requests() generates 200 requests spread across 40 users, which gives us enough traffic to see meaningful differences between strategies while keeping the simulation lightweight. Copy CodeCopiedUse a different Browser import random import hashlib random.seed(42) def legacy_model(request): return {“model”: “legacy”, “score”: random.random() * 0.35} def candidate_model(request): return {“model”: “candidate”, “score”: random.random() * 0.55} def make_requests(n=200): users = [f”user_{i}” for i in range(40)] return [{“id”: f”req_{i}”, “user”: random.choice(users)} for i in range(n)] requests = make_requests() A/B Testing ab_route() is the core of this strategy — for every incoming request, it draws a random number and routes to the candidate model only if that number falls below 0.10, otherwise the request goes to legacy. This gives the candidate roughly 10% of traffic. We then collect the prediction scores from each model separately and compute the average at the end. In a real system, these scores would be replaced by actual engagement metrics like click-through rate or watch time — here the score just stands in for “how good was this recommendation.” Copy CodeCopiedUse a different Browser print(“── 1. A/B Testing ──────────────────────────────────────────”) CANDIDATE_TRAFFIC = 0.10 # 10 % of requests go to candidate def ab_route(request): return candidate_model if random.random() < CANDIDATE_TRAFFIC else legacy_model results = {“legacy”: [], “candidate”: []} for req in requests: model = ab_route(req) pred = model(req) results[pred[“model”]].append(pred[“score”]) for name, scores in results.items(): print(f” {name:12s} | requests: {len(scores):3d} | avg score: {sum(scores)/len(scores):.3f}”) Canary Testing The key function here is get_canary_users(), which uses an MD5 hash to deterministically assign users to the canary group. The important word is deterministic — sorting users by their hash means the

Safely Deploying ML Models to Production: Four Controlled Strategies (A/B, Canary, Interleaved, Shadow Testing) Leer entrada »

AI, Committee, Noticias, Uncategorized

A Coding Implementation for Building and Analyzing Crystal Structures Using Pymatgen for Symmetry Analysis, Phase Diagrams, Surface Generation, and Materials Project Integration

In this tutorial, we explore the capabilities of the pymatgen library for computational materials science using Python. We begin by constructing crystal structures such as silicon, sodium chloride, and a LiFePO₄-like material, and then investigate their lattice properties, densities, and compositions. Also, we analyze symmetry using space-group detection, examine atomic coordination environments, and apply oxidation-state decorations to better understand the structures’ chemistry. We also generate supercells, perturb atomic positions, and compute distance matrices to study structural relationships at larger scales. Along the way, we simulate X-ray diffraction patterns, construct a simple phase diagram, and demonstrate how disordered alloy structures can be approximated by ordered configurations. Finally, we extend the workflow to include molecule analysis, CIF export, and optional querying of the Materials Project database, thereby illustrating how pymatgen can serve as a powerful toolkit for materials modeling and data analysis. Copy CodeCopiedUse a different Browser !pip -q install pymatgen mp-api spglib import os import json import warnings import sys warnings.filterwarnings(“ignore”) import numpy as np import pandas as pd import matplotlib.pyplot as plt from pymatgen.core import Lattice, Structure, Molecule from pymatgen.core.surface import SlabGenerator from pymatgen.core.composition import Composition from pymatgen.symmetry.analyzer import SpacegroupAnalyzer from pymatgen.analysis.local_env import CrystalNN from pymatgen.analysis.diffraction.xrd import XRDCalculator from pymatgen.analysis.phase_diagram import PDEntry, PhaseDiagram from pymatgen.transformations.standard_transformations import ( SupercellTransformation, OrderDisorderedStructureTransformation, OxidationStateDecorationTransformation, ) from pymatgen.io.cif import CifWriter print(“Python:”, sys.version.split()[0]) print(“NumPy:”, np.__version__) print(“pandas:”, pd.__version__) try: import pymatgen print(“pymatgen:”, pymatgen.__version__) except Exception: import importlib.metadata print(“pymatgen:”, importlib.metadata.version(“pymatgen”)) def line(): print(“=” * 100) def header(title): line() print(title) line() header(“1. BUILD EXAMPLE STRUCTURES”) si = Structure( Lattice.cubic(5.431), [“Si”, “Si”], [[0, 0, 0], [0.25, 0.25, 0.25]], ) nacl = Structure( Lattice.cubic(5.64), [“Na”, “Cl”], [[0, 0, 0], [0.5, 0.5, 0.5]], ) li_fe_po4 = Structure( Lattice.orthorhombic(10.33, 6.01, 4.69), [“Li”, “Fe”, “P”, “O”, “O”, “O”, “O”], [ [0.0, 0.0, 0.0], [0.5, 0.5, 0.5], [0.1, 0.25, 0.2], [0.22, 0.04, 0.28], [0.72, 0.54, 0.78], [0.31, 0.66, 0.12], [0.81, 0.16, 0.62], ], ) for name, s in [(“Si”, si), (“NaCl”, nacl), (“LiFePO4-like”, li_fe_po4)]: print(f”{name}: formula={s.composition.formula}, sites={len(s)}, volume={s.volume:.3f} Å^3″) We begin by installing the required libraries. We initialize the environment, verify package versions, and define helper functions to organize the output. We then construct example crystal structures such as silicon, NaCl, and a LiFePO₄-like structure and print their basic structural properties. Copy CodeCopiedUse a different Browser header(“2. BASIC INTROSPECTION”) for name, s in [(“Si”, si), (“NaCl”, nacl), (“LiFePO4-like”, li_fe_po4)]: print(f”n{name}”) print(“Reduced formula:”, s.composition.reduced_formula) print(“Density:”, round(s.density, 4), “g/cm^3”) print(“Lattice parameters (a, b, c):”, tuple(round(x, 4) for x in s.lattice.abc)) print(“Angles (alpha, beta, gamma):”, tuple(round(x, 4) for x in s.lattice.angles)) print(“First site:”, s[0]) header(“3. SPACE GROUP AND SYMMETRY ANALYSIS”) for name, s in [(“Si”, si), (“NaCl”, nacl), (“LiFePO4-like”, li_fe_po4)]: sga = SpacegroupAnalyzer(s, symprec=0.1) print(f”n{name}”) print(“Space group symbol:”, sga.get_space_group_symbol()) print(“Space group number:”, sga.get_space_group_number()) print(“Crystal system:”, sga.get_crystal_system()) print(“Lattice type:”, sga.get_lattice_type()) print(“Primitive sites:”, len(sga.find_primitive())) print(“Conventional sites:”, len(sga.get_conventional_standard_structure())) We examine the structures in greater detail by inspecting their formulas, densities, lattice parameters, and site information. We then perform a symmetry analysis using SpacegroupAnalyzer to determine space-group symbols, crystal systems, and lattice types. Through this step, we gain insight into the crystallographic symmetry and structural characteristics of the materials. Copy CodeCopiedUse a different Browser header(“4. LOCAL ENVIRONMENT WITH CRYSTALNN”) cnn = CrystalNN() def summarize_neighbors(structure, label): print(f”n{label}”) for i, site in enumerate(structure[:min(4, len(structure))]): try: nn_info = cnn.get_nn_info(structure, i) species = [str(x[“site”].specie) for x in nn_info] weights = [round(float(x[“weight”]), 3) for x in nn_info] print(f”Site {i} {site.species_string}: CN={len(nn_info)}, neighbors={species}, weights={weights}”) except Exception as e: print(f”Site {i} {site.species_string}: neighbor analysis failed -> {e}”) summarize_neighbors(si, “Si”) summarize_neighbors(nacl, “NaCl”) header(“5. OXIDATION STATE DECORATION”) oxi_transform = OxidationStateDecorationTransformation( {“Li”: 1, “Fe”: 2, “P”: 5, “O”: -2, “Na”: 1, “Cl”: -1, “Si”: 0} ) nacl_oxi = oxi_transform.apply_transformation(nacl.copy()) lfp_oxi = oxi_transform.apply_transformation(li_fe_po4.copy()) print(“NaCl species with oxidation states:”, [str(site.specie) for site in nacl_oxi]) print(“LiFePO4-like species with oxidation states:”, [str(site.specie) for site in lfp_oxi]) We analyze the local atomic environments using the CrystalNN coordination analysis algorithm. We identify neighboring atoms for selected sites and evaluate their coordination numbers and weights. We then decorate the structures with oxidation states to better represent the chemical environment. Copy CodeCopiedUse a different Browser header(“6. MAKE SUPERCELLS”) si_super = SupercellTransformation([[2, 0, 0], [0, 2, 0], [0, 0, 2]]).apply_transformation(si.copy()) nacl_super = SupercellTransformation([[2, 0, 0], [0, 2, 0], [0, 0, 2]]).apply_transformation(nacl.copy()) print(“Si supercell sites:”, len(si_super), “formula:”, si_super.composition.formula) print(“NaCl supercell sites:”, len(nacl_super), “formula:”, nacl_super.composition.formula) header(“7. PERTURB STRUCTURE AND COMPUTE DISTANCE MATRIX”) si_perturbed = si_super.copy() si_perturbed.translate_sites([0], [0.01, -0.005, 0.012], frac_coords=False) dm = si_perturbed.distance_matrix print(“Distance matrix shape:”, dm.shape) print(“First 5 distances from site 0:”, np.round(dm[0][:5], 4)) header(“8. GENERATE A SURFACE SLAB”) slabgen = SlabGenerator( initial_structure=si, miller_index=(1, 1, 1), min_slab_size=8.0, min_vacuum_size=12.0, center_slab=True, in_unit_planes=False, ) slabs = slabgen.get_slabs() slab = slabs[0] print(“Number of generated slabs:”, len(slabs)) print(“Chosen slab formula:”, slab.composition.formula) print(“Chosen slab sites:”, len(slab)) print(“Chosen slab lattice:”, tuple(round(x, 3) for x in slab.lattice.abc)) We expand the crystal structures into larger supercells to study periodic structures at a larger scale. We apply a small perturbation to atomic positions and compute the resulting distance matrix to analyze structural changes. We also generate a surface slab from the silicon crystal to demonstrate how surface structures can be modeled. Copy CodeCopiedUse a different Browser header(“9. XRD SIMULATION”) xrd = XRDCalculator(wavelength=”CuKa”) pattern_si = xrd.get_pattern(si, two_theta_range=(10, 90)) pattern_nacl = xrd.get_pattern(nacl, two_theta_range=(10, 90)) plt.figure(figsize=(12, 4)) plt.vlines(pattern_si.x, [0], pattern_si.y, linewidth=1.5) plt.xlabel(r”2$theta$ (degrees)”) plt.ylabel(“Intensity”) plt.title(“Simulated XRD Pattern: Si”) plt.show() plt.figure(figsize=(12, 4)) plt.vlines(pattern_nacl.x, [0], pattern_nacl.y, linewidth=1.5) plt.xlabel(r”2$theta$ (degrees)”) plt.ylabel(“Intensity”) plt.title(“Simulated XRD Pattern: NaCl”) plt.show() header(“10. SIMPLE PHASE DIAGRAM”) entries = [ PDEntry(Composition(“Li”), 0.0), PDEntry(Composition(“Fe”), 0.0), PDEntry(Composition(“P”), 0.0), PDEntry(Composition(“O2”), 0.0), PDEntry(Composition(“Li2O”), -6.0), PDEntry(Composition(“FeO”), -4.2), PDEntry(Composition(“Fe2O3”), -10.5), PDEntry(Composition(“P2O5”), -15.0), PDEntry(Composition(“Li3PO4”), -18.5), PDEntry(Composition(“FePO4”), -12.2), PDEntry(Composition(“LiFePO4”), -16.9), ] pdg = PhaseDiagram(entries) target = [e for e in entries if e.composition.reduced_formula == “LiFePO4”][0] e_above_hull = pdg.get_e_above_hull(target) decomp, e_hull = pdg.get_decomp_and_e_above_hull(target) print(“Target entry:”, target.composition.reduced_formula) print(“Energy above hull:”, round(float(e_above_hull), 6), “eV/atom”) print(“Decomposition products:”) for k, v in decomp.items(): print(” “, k.composition.reduced_formula, “:”, round(float(v), 6)) We simulate X-ray diffraction patterns for silicon and NaCl using pymatgen’s diffraction tools. We visualize the diffraction peaks to understand how the crystal structure influences the XRD pattern. We then construct a simple thermodynamic phase diagram and calculate the stability of LiFePO₄ relative to competing phases. Copy CodeCopiedUse

A Coding Implementation for Building and Analyzing Crystal Structures Using Pymatgen for Symmetry Analysis, Phase Diagrams, Surface Generation, and Materials Project Integration Leer entrada »

AI, Committee, Noticias, Uncategorized

OpenAI is throwing everything into building a fully automated researcher

OpenAI is refocusing its research efforts and throwing its resources into a new grand challenge. The San Francisco firm has set its sights on building what it calls an AI researcher, a fully automated agent-based system that will be able to go off and tackle large, complex problems by itself. ​​OpenAI says that this new research goal will be its “North Star” for the next few years, pulling together multiple research strands, including work on reasoning models, agents, and interpretability. There’s even a timeline. OpenAI plans to build “an autonomous AI research intern”—a system that can take on a small number of specific research problems by itself—by September. The AI intern will be the precursor to a fully automated multi-agent research system that the company plans to debut in 2028. This AI researcher (OpenAI says) will be able to tackle problems that are too large or complex for humans to cope with. Those tasks might be related to math and physics—such as coming up with new proofs or conjectures—or life sciences like biology and chemistry, or even business and policy dilemmas. In theory, you would throw such a tool any kind of problem that can be formulated in text, code, or whiteboard scribbles—which covers a lot. OpenAI has been setting the agenda for the AI industry for years. Its early dominance with large language models shaped the technology that hundreds of millions of people use every day. But it now faces fierce competition from rival model makers like Anthropic and Google DeepMind. What OpenAI decides to build next matters—for itself and for the future of AI.    A big part of that decision falls to Jakub Pachocki, OpenAI’s chief scientist, who sets the company’s long-term research goals. Pachocki played key roles in the development of both GPT-4, a game-changing LLM released in 2023, and so-called reasoning models, a technology that first appeared in 2024 and now underpins all major chatbots and agent-based systems.  In an exclusive interview this week, Pachocki talked me through OpenAI’s latest vision. “I think we are getting close to a point where we’ll have models capable of working indefinitely in a coherent way just like people do,” he says. “Of course, you still want people in charge and setting the goals. But I think we will get to a point where you kind of have a whole research lab in a data center.” Solving hard problems Such big claims aren’t new. Saving the world by solving its hardest problems is the stated mission of all the top AI firms. Demis Hassabis told me back in 2022 that it was why he started DeepMind. Anthropic CEO Dario Amodei says he is building the equivalent of a country of geniuses in a data center. Pachocki’s boss, Sam Altman, wants to cure cancer. But Pachocki says OpenAI now has most of what it needs to get there. In January, OpenAI released Codex, an agent-based app that can spin up code on the fly to carry out tasks on your computer. It can analyze documents, generate charts, make you a daily digest of your inbox and social media, and much more. (Other firms have released similar tools, such as Anthropic’s Claude Code and Claude Cowork.) OpenAI claims that most of its technical staffers now use Codex in their work. You can look at Codex as a very early version of the AI researcher, says Pachocki: “I expect Codex to get fundamentally better.” The key is to make a system that can run for longer periods of time, with less human guidance. “What we’re really looking at for an automated research intern is a system that you can delegate tasks [to] that would take a person a few days,” says Pachocki. “There are a lot of people excited about building systems that can do more long-running scientific research,” says Doug Downey, a research scientist at the Allen Institute for AI, who is not connected to OpenAI. “I think it’s largely driven by the success of these coding agents. The fact that you can delegate quite substantial coding tasks to tools like Codex is incredibly useful and incredibly impressive. And it raises the question: Can we do similar things outside coding, in broader areas of science?” For Pachocki, that’s a clear Yes. In fact, he thinks it’s just a matter of pushing ahead on the path we’re already on. A simple boost in all-round capability also leads to models that can work longer without help, he says. He points to the leap from 2020’s GPT-3 to 2023’s GPT-4, two of OpenAI’s previous models. GPT-4 was able to work on a problem for far longer than its predecessor, even without specialized training, he says.  So-called reasoning models brought another bump. Training LLMs to work through problems step by step, backtracking when they make a mistake or hit a dead end, has also made models better at working for longer periods of time. And Pachocki is convinced that OpenAI’s reasoning models will continue to get better. But OpenAI is also training its systems to work by themselves for longer by feeding them specific samples of complex tasks, such as hard puzzles taken from math and coding contests, which force the models to learn how to do things like keep track of very large chunks of text and split problems up into (and then manage) multiple subtasks. The aim isn’t to build models that just win math competitions. “That lets you prove that the technology works before you connect it to the real world,” says Pachocki. “If we really wanted to, we could build an amazing automated mathematician. We have all the tools, and I think it would be relatively easy. But it’s not something we’re going to prioritize now because, you know, at the point where you believe you can do it, there’s much more urgent things to do.” “We are much more focused now on research that’s relevant in the real world,” he adds. Right now that means taking what Codex can do

OpenAI is throwing everything into building a fully automated researcher Leer entrada »

We use cookies to improve your experience and performance on our website. You can learn more at Política de privacidad and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
es_ES