YouZum

ニュース

AI, Committee, ニュース, Uncategorized

BOND 2025 AI Trends Report Shows AI Ecosystem Growing Faster than Ever with Explosive User and Developer Adoption

BOND’s latest report on Trends – Artificial Intelligence (May 2025) presents a comprehensive data-driven snapshot of the current state and rapid evolution of AI technology. The report highlights some striking trends underscoring the unprecedented velocity of AI adoption, technological improvement, and market impact. This article reviews several key findings from the report and explores their implications for the AI ecosystem. Explosive Adoption of Open-Source Large Language Models One of the standout observations is the remarkable uptake of Meta’s Llama models. Over an eight-month span, Llama downloads surged by a factor of 3.4×, marking an unprecedented developer adoption curve for any open-source large language model (LLM). This acceleration highlights the expanding democratization of AI capabilities beyond proprietary platforms, enabling a broad spectrum of developers to integrate and innovate with advanced models. Source: https://www.bondcap.com/reports/tai The rapid acceptance of Llama illustrates a growing trend in the industry: open-source AI projects are becoming competitive alternatives to proprietary models, fueling a more distributed ecosystem. This proliferation accelerates innovation cycles and lowers barriers to entry for startups and research groups. AI Chatbots Achieving Human-Level Conversational Realism The report also documents significant advances in conversational AI. In Q1 2025, Turing-style tests showed that human evaluators mistook AI chatbot responses for human replies 73% of the time—a substantial jump from approximately 50% only six months prior. This rapid improvement reflects the growing sophistication of LLMs in mimicking human conversational nuances such as context retention, emotional resonance, and colloquial expression. Source: https://www.bondcap.com/reports/tai This trend has profound implications for industries reliant on customer interaction, including support, sales, and personal assistants. As chatbots approach indistinguishability from humans in conversation, businesses will need to rethink user experience design, ethical considerations, and transparency standards to maintain trust. ChatGPT’s Search Volume Surpasses Google’s Early Growth by 5.5× ChatGPT reached an estimated 365 billion annual searches within just two years of its public launch in November 2022. This growth rate outpaces Google’s trajectory, which took 11 years (1998–2009) to reach the same volume of annual searches. In essence, ChatGPT’s search volume ramped up about 5.5 times faster than Google’s did. Source: https://www.bondcap.com/reports/tai This comparison underscores the transformative shift in how users interact with information retrieval systems. The conversational and generative nature of ChatGPT has fundamentally altered expectations for search and discovery, accelerating adoption and daily engagement. NVIDIA’s GPUs Power Massive AI Throughput Gains While Reducing Power Draw Between 2016 and 2024, NVIDIA GPUs achieved a 225× increase in AI inference throughput, while simultaneously cutting data center power consumption by 43%. This impressive dual improvement has yielded an astounding >30,000× increase in theoretical annual token processing capacity per $1 billion data center investment. Source: https://www.bondcap.com/reports/tai This leap in efficiency underpins the scalability of AI workloads and dramatically lowers the operational cost of AI deployments. As a result, enterprises can now deploy larger, more complex AI models at scale with reduced environmental impact and better cost-effectiveness. DeepSeek’s Rapid User Growth Captures a Third of China’s Mobile AI Market In the span of just four months, from January to April 2025, DeepSeek scaled from zero to 54 million monthly active mobile AI users in China, securing over 34% market share in the mobile AI segment. This rapid growth reflects both the enormous demand in China’s mobile AI ecosystem and DeepSeek’s ability to capitalize on it through local market understanding and product fit. Source: https://www.bondcap.com/reports/tai The speed and scale of DeepSeek’s adoption also highlight the growing global competition in AI innovation, particularly between China and the U.S., with localized ecosystems developing rapidly in parallel. The Revenue Opportunity for AI Inference Has Skyrocketed The report outlines a massive shift in the potential revenue from AI inference tokens processed in large data centers. In 2016, a $1 billion-scale data center could process roughly 5 trillion inference tokens annually, generating about $24 million in token-related revenue. By 2024, that same investment could handle an estimated 1,375 trillion tokens per year, translating to nearly $7 billion in theoretical revenue — a 30,000× increase. Source: https://www.bondcap.com/reports/tai This enormous leap stems from improvements in both hardware efficiency and algorithmic optimizations that dramatically reduce inference costs. The Plunge in AI Inference Costs One of the key enablers of these trends is the steep decline in inference costs per million tokens. For example, the cost to generate a million tokens using GPT-3.5 dropped from over $10 in September 2022 to around $1 by mid-2023. ChatGPT’s cost per 75-word response approached near zero within its first year. This precipitous fall in pricing closely mirrors historical cost declines in other technologies, such as computer memory, which fell to near zero over two decades, and electric power, which dropped to about 2–3% of its initial price after 60–70 years. In contrast, more static costs like that of light bulbs have remained largely flat over time. The IT Consumer Price Index vs. Compute Demand BOND’s report also examines the relationship between IT consumer price trends and compute demand. Since 2010, compute requirements for AI have increased by approximately 360% per year, leading to an estimated total of 10²⁶ floating point operations (FLOPs) in 2024. During the same period, the IT consumer price index fell from 100 to below 10, indicating dramatically cheaper hardware costs. This decoupling means organizations can train larger and more complex AI models while spending significantly less on compute infrastructure, further accelerating AI innovation cycles. Conclusion BOND’s Trends – Artificial Intelligence report offers compelling quantitative evidence that AI is evolving at an unprecedented pace. The combination of rapid user adoption, explosive developer engagement, hardware efficiency breakthroughs, and falling inference costs is reshaping the AI landscape globally. From Meta’s Llama open-source surge to DeepSeek’s rapid market capture in China, and from ChatGPT’s hyper-accelerated search growth to NVIDIA’s remarkable GPU performance gains, the data reflect a highly dynamic ecosystem. The steep decline in AI inference costs amplifies this effect, enabling new applications and business models. The key takeaway for AI practitioners and industry watchers is clear: AI’s technological and economic momentum is accelerating, demanding continuous innovation and strategic agility.

BOND 2025 AI Trends Report Shows AI Ecosystem Growing Faster than Ever with Explosive User and Developer Adoption 投稿を読む »

AI, Committee, ニュース, Uncategorized

Yandex Releases Yambda: The World’s Largest Event Dataset to Accelerate Recommender Systems

Yandex has recently made a significant contribution to the recommender systems community by releasing Yambda, the world’s largest publicly available dataset for recommender system research and development. This dataset is designed to bridge the gap between academic research and industry-scale applications, offering nearly 5 billion anonymized user interaction events from Yandex Music — one of the company’s flagship streaming services with over 28 million monthly users. Why Yambda Matters: Addressing a Critical Data Gap in Recommender Systems Recommender systems underpin the personalized experiences of many digital services today, from e-commerce and social networks to streaming platforms. These systems rely heavily on massive volumes of behavioral data, such as clicks, likes, and listens, to infer user preferences and deliver tailored content. However, the field of recommender systems has lagged behind other AI domains, like natural language processing, largely due to the scarcity of large, openly accessible datasets. Unlike large language models (LLMs), which learn from publicly available text sources, recommender systems need sensitive behavioral data — which is commercially valuable and hard to anonymize. As a result, companies have traditionally guarded this data closely, limiting researchers’ access to real-world-scale datasets. Existing datasets such as Spotify’s Million Playlist Dataset, Netflix Prize data, and Criteo’s click logs are either too small, lack temporal detail, or are poorly documented for developing production-grade recommender models. Yandex’s release of Yambda addresses these challenges by providing a high-quality, extensive dataset with a rich set of features and anonymization safeguards. What Yambda Contains: Scale, Richness, and Privacy The Yambda dataset comprises 4.79 billion anonymized user interactions collected over a 10-month period. These events come from roughly 1 million users interacting with nearly 9.4 million tracks on Yandex Music. The dataset includes: User Interactions: Both implicit feedback (listens) and explicit feedback (likes, dislikes, and their removals). Anonymized Audio Embeddings: Vector representations of tracks derived from convolutional neural networks, enabling models to leverage audio content similarity. Organic Interaction Flags: An “is_organic” flag indicates whether users discovered a track independently or via recommendations, facilitating behavioral analysis. Precise Timestamps: Each event is timestamped to preserve temporal ordering, crucial for modeling sequential user behavior. All user and track identifiers are anonymized using numeric IDs to comply with privacy standards, ensuring no personally identifiable information is exposed. The dataset is provided in Apache Parquet format, which is optimized for big data processing frameworks like Apache Spark and Hadoop, and also compatible with analytical libraries such as Pandas and Polars. This makes Yambda accessible for researchers and developers working in diverse environments. Evaluation Method: Global Temporal Split A key innovation in Yandex’s dataset is the adoption of a Global Temporal Split (GTS) evaluation strategy. In typical recommender system research, the widely used Leave-One-Out method removes the last interaction of each user for testing. However, this approach disrupts the temporal continuity of user interactions, creating unrealistic training conditions. GTS, on the other hand, splits the data based on timestamps, preserving the entire sequence of events. This approach mimics real-world recommendation scenarios more closely because it prevents any future data from leaking into training and allows models to be tested on truly unseen, chronologically later interactions. This temporal-aware evaluation is essential for benchmarking algorithms under realistic constraints and understanding their practical effectiveness. Baseline Models and Metrics Included To support benchmarking and accelerate innovation, Yandex provides baseline recommender models implemented on the dataset, including: MostPop: A popularity-based model recommending the most popular items. DecayPop: A time-decayed popularity model. ItemKNN: A neighborhood-based collaborative filtering method. iALS: Implicit Alternating Least Squares matrix factorization. BPR: Bayesian Personalized Ranking, a pairwise ranking method. SANSA and SASRec: Sequence-aware models leveraging self-attention mechanisms. These baselines are evaluated using standard recommender metrics such as: NDCG@k (Normalized Discounted Cumulative Gain): Measures ranking quality emphasizing the position of relevant items. Recall@k: Assesses the fraction of relevant items retrieved. Coverage@k: Indicates the diversity of recommendations across the catalog. Providing these benchmarks helps researchers quickly gauge the performance of new algorithms relative to established methods. Broad Applicability Beyond Music Streaming While the dataset originates from a music streaming service, its value extends far beyond that domain. The interaction types, user behavior dynamics, and large scale make Yambda a universal benchmark for recommender systems across sectors like e-commerce, video platforms, and social networks. Algorithms validated on this dataset can be generalized or adapted to various recommendation tasks. Benefits for Different Stakeholders Academia: Enables rigorous testing of theories and new algorithms at an industry-relevant scale. Startups and SMBs: Offers a resource comparable to what tech giants possess, leveling the playing field and accelerating the development of advanced recommendation engines. End Users: Indirectly benefits from smarter recommendation algorithms that improve content discovery, reduce search time, and increase engagement. My Wave: Yandex’s Personalized Recommender System Yandex Music leverages a proprietary recommender system called My Wave, which incorporates deep neural networks and AI to personalize music suggestions. My Wave analyzes thousands of factors including: User interaction sequences and listening history. Customizable preferences such as mood and language. Real-time music analysis of spectrograms, rhythm, vocal tone, frequency ranges, and genres. This system dynamically adapts to individual tastes by identifying audio similarities and predicting preferences, demonstrating the kind of complex recommendation pipeline that benefits from large-scale datasets like Yambda. Ensuring Privacy and Ethical Use The release of Yambda underscores the importance of privacy in recommender system research. Yandex anonymizes all data with numeric IDs and omits personally identifiable information. The dataset contains only interaction signals without revealing exact user identities or sensitive attributes. This balance between openness and privacy allows for robust research while protecting individual user data, a critical consideration for the ethical advancement of AI technologies. Access and Versions Yandex offers the Yambda dataset in three sizes to accommodate different research and computational capacities: Full version: ~5 billion events. Medium version: ~500 million events. Small version: ~50 million events. All versions are accessible via Hugging Face, a popular platform for hosting datasets and machine learning models, enabling easy integration into research workflows. Conclusion Yandex’s release of the Yambda dataset marks a pivotal moment

Yandex Releases Yambda: The World’s Largest Event Dataset to Accelerate Recommender Systems 投稿を読む »

AI, Committee, ニュース, Uncategorized

Multimodal Foundation Models Fall Short on Physical Reasoning: PHYX Benchmark Highlights Key Limitations in Visual and Symbolic Integration

State-of-the-art models show human-competitive accuracy on AIME, GPQA, MATH-500, and OlympiadBench, solving Olympiad-level problems. Recent multimodal foundation models have advanced benchmarks for disciplinary knowledge and mathematical reasoning. However, these evaluations miss a crucial aspect of machine intelligence: physical reasoning, which requires integrating disciplinary knowledge, symbolic operations, and real-world constraints. Physical problem-solving differs fundamentally from pure mathematical reasoning as it demands models to decode implicit conditions in questions. For example, interpreting “smooth surface” as zero friction coefficient, and maintaining physical consistency across reasoning chains because physical laws remain constant regardless of reasoning trajectories. MLLM shows excellent visual understanding by integrating visual and textual data across various tasks, motivating exploration of its reasoning abilities. However, uncertainty remains regarding whether these models possess genuine advanced reasoning capabilities for visual tasks, particularly in physical domains closer to real-world scenarios. Several LLM benchmarks have emerged to evaluate reasoning abilities, with PHYBench being most relevant for physics reasoning. MLLM scientific benchmarks, such as PhysReason and EMMA, contain multimodal physics problems with figures, however, they include only small physics subsets, which inadequately evaluate MLLMs’ capabilities for reasoning and solving advanced physics problems. Researchers from the University of Hong Kong, the University of Michigan, the University of Toronto, the University of Waterloo, and the Ohio State University have proposed PHYX, a novel benchmark to evaluate the physical reasoning capabilities of foundation models. It comprises 3,000 visually-grounded physics questions, precisely curated across six distinct physics domains: Mechanics, Electromagnetism, Thermodynamics, Wave/Acoustics, Optics, and Modern Physics. It evaluates physics-based reasoning via multimodal problem-solving with three core innovations: (a) 3,000 newly collected questions with realistic physical scenarios requiring integrated visual analysis and causal reasoning, (b) Expert-validated data design covering six fundamental physics domains, and (c) Strict unified three-step evaluation protocols. Researchers designed a four-stage data collection process to ensure high-quality data. The process begins with an in-depth survey of core physics disciplines to determine coverage across diverse domains and subfields, followed by the recruitment of STEM graduate students as expert annotators. They comply with copyright restrictions and avoid data contamination by selecting questions without answers that are immediately available. Moreover, quality control involves a three-stage cleaning process including duplicate detection through lexical overlap analysis with manual review by physics Ph.D. students, followed by filtering the shortest 10% of questions based on textual length, resulting in 3,000 high-quality questions from an initial collection of 3,300. PHYX presents significant challenges for current models, with even the worst-performing human experts achieving 75.6% accuracy, outperforming all evaluated models and showing a gap between human expertise and current model capabilities. The benchmark reveals that multiple-choice formats narrow performance gaps by allowing weaker models to rely on surface-level cues, but open-ended questions demand genuine reasoning and precise answer generation. Comparing GPT-4o’s performance on PHYX to previously reported results on MathVista and MATH-V (both 63.8%), lower accuracy in physical reasoning tasks emphasizes that physical reasoning requires deeper integration of abstract concepts and real-world knowledge, presenting greater challenges than purely mathematical contexts. In conclusion, researchers introduced PHYX, the first large-scale benchmark for evaluating physical reasoning in multimodal, visually grounded scenarios. Rigorous evaluation reveals that state-of-the-art models show limitations in physical reasoning, relying predominantly on memorized knowledge, mathematical formulas, and superficial visual patterns rather than genuine understanding of physical principles. The benchmark focuses exclusively on English-language prompts and annotations, limiting assessment of multilingual reasoning abilities. Also, while images depict physically realistic scenarios, they are often schematic or textbook-style rather than real-world photographs, which may not fully capture the complexity of perception in natural environments. Check out the Paper, Code and Project Page. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 95k+ ML SubReddit and Subscribe to our Newsletter. The post Multimodal Foundation Models Fall Short on Physical Reasoning: PHYX Benchmark Highlights Key Limitations in Visual and Symbolic Integration appeared first on MarkTechPost.

Multimodal Foundation Models Fall Short on Physical Reasoning: PHYX Benchmark Highlights Key Limitations in Visual and Symbolic Integration 投稿を読む »

AI, Committee, ニュース, Uncategorized

Apple and Duke Researchers Present a Reinforcement Learning Approach That Enables LLMs to Provide Intermediate Answers, Enhancing Speed and Accuracy

Long CoT reasoning improves large language models’ performance on complex tasks but comes with drawbacks. The typical “think-then-answer” method slows down response times, disrupting real-time interactions like those in chatbots. It also risks inaccuracies, as errors in earlier reasoning steps can lead to a misleading final answer. Unlike humans, who often share partial thoughts or conclusions during conversations, LLMs delay responses until all reasoning is complete. While RL is commonly used to train reasoning models, it mainly rewards final answers, overlooking useful intermediate insights. There is growing interest in teaching models that alternate between thinking and answering, but this remains a challenge.  RL has become a popular method to enhance reasoning in LLMs, building on its success in aligning models with human preferences. Two common reward types guide RL: outcome-based rewards (ORM), which focus on the final answer, and process-based rewards (PRM), which provide feedback on intermediate reasoning steps. While PRMs offer more detailed supervision, they often rely on human annotation and additional models, making them complex and prone to issues like reward hacking. Separately, efforts to improve LLM reasoning have explored prompting strategies, structured reasoning, tool integration, and methods to reduce latency and improve efficiency.  Researchers from Apple and Duke University introduce Interleaved Reasoning, a new RL approach that enables language models to alternate between thinking and answering when solving complex, multi-step questions. Instead of waiting until the end to respond, models provide informative intermediate answers, which improves feedback for users and guides their reasoning. Using a straightforward rule-based reward, the model is trained to produce helpful reasoning steps, leading to over 80% faster responses and up to 19.3% better accuracy. Trained only on QA and logic datasets, the method demonstrates strong generalization to more challenging benchmarks, such as MATH, GPQA, and MMLU.  The study proposes a reinforcement learning framework to train LLMs for Interleaved Reasoning, where models alternate between internal thinking and user-facing intermediate answers. Each intermediate step, or “sub-answer,” is shared once the model reaches a meaningful milestone in reasoning. A specialized training template with <think> and <answer> tags is used. The approach utilizes rule-based rewards—specifically, format, final accuracy, and conditional intermediate accuracy—to guide learning. Notably, intermediate rewards are applied only when specific criteria are met, ensuring the model prioritizes overall correctness. They also test different reward schemes, such as all-or-none, partial credit, and time-discounted rewards, to optimize the quality of reasoning.  The interleaved reasoning approach was evaluated on both familiar and unfamiliar datasets using Qwen2.5 models (1.5B and 7B). Unlike traditional methods that separate thinking and answering, the interleaved method provides answers incrementally, improving both speed and usefulness. When combined with intermediate rewards, it significantly enhances model performance while reducing response delays by over 80%. Even without exposure to new domains during training, the model adapts well, showing strong generalization. These results highlight the value of interleaved reasoning in making AI systems more responsive and effective in real-world, multi-step reasoning tasks.  In conclusion, the study explores how interleaved reasoning—where models alternate between reasoning and generating intermediate answers—can significantly improve performance and responsiveness. Using the Qwen2.5-1.5B model, the authors show that providing timely intermediate feedback during training boosts accuracy and accelerates response generation. Different RL strategies were tested, with PPO showing stable results, and conditional, time-discounted rewards proving to be the most effective. The method scales well to complex tasks and outperforms traditional think-then-answer baselines. Unlike token-level reward models, this approach employs simple rule-based rewards after completing full reasoning steps, thereby avoiding reward hacking. Ultimately, interleaved reasoning enhances reasoning quality and efficiency without relying on external tools.  Check out the Paper. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 95k+ ML SubReddit and Subscribe to our Newsletter. The post Apple and Duke Researchers Present a Reinforcement Learning Approach That Enables LLMs to Provide Intermediate Answers, Enhancing Speed and Accuracy appeared first on MarkTechPost.

Apple and Duke Researchers Present a Reinforcement Learning Approach That Enables LLMs to Provide Intermediate Answers, Enhancing Speed and Accuracy 投稿を読む »

ja