Hardware & InfraBusiness & FundingSoftBank says it will invest up to €75 billion to build French data centersSoftBank's €75 billion commitment to French data center infrastructure signals accelerating capital deployment in European AI compute. The planned 5 gigawatt expansion addresses a critical bottleneck for large-scale model training and inference across the continent, positioning France as a strategic hub amid US-China chip competition and EU sovereignty concerns. This move reflects how major tech investors now treat data center buildout as foundational to AI dominance, with implications for model availability, latency, and geopolitical leverage in the AI supply chain.TechCrunch - AI·May 3081
ResearchProducts & AppsHow we contain Claude across productsAnthropic published detailed technical documentation on how it isolates Claude across multiple deployment surfaces, including process sandboxes, virtual machines, filesystem restrictions, and network egress controls. The move addresses a critical gap in the AI industry: most sandbox implementations remain opaque, making it difficult for users and enterprises to assess genuine containment guarantees. By transparently explaining the layered constraints that prevent agents from exceeding their intended scope, Anthropic sets a precedent for security disclosure that could reshape how the field approaches agent safety and user trust in production systems.Simon Willison·May 3084
ResearchCitation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation GraphsResearchers have developed citation grounding, a systematic framework to measure and mitigate hallucinations in legal LLM outputs by validating generated citations against a ground-truth graph of 100.8 million Ukrainian court decisions. The approach decomposes hallucination into three diagnostic categories: existence verification, contextual relevance, and temporal validity. This work addresses a critical failure mode in high-stakes domains where fabricated or outdated legal references carry real consequences, establishing a replicable methodology that could extend beyond law to other citation-dependent fields like medicine and academia.arXiv cs.CL·May 3062
ResearchTools & CodeChunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and LimitationsA systematic evaluation of chunking strategies in RAG systems addresses a critical gap in LLM infrastructure. While fixed-size and semantic chunking dominate production systems, emerging methods proliferate with narrow validation and unclear trade-offs between retrieval quality and computational overhead. This first comparative study matters because chunking directly impacts both retrieval accuracy and inference cost, yet practitioners lack principled guidance on method selection across diverse data types and use cases. The findings will shape how teams architect retrieval pipelines at scale.arXiv cs.CL·May 3062
Opinion & AnalysisI Am Retiring from Tech to Live OfflineChad Whitacre, a prominent open-source developer, is abandoning tech entirely in response to AI's trajectory, citing it as the breaking point after years of industry strain. His typewritten, deliberately analog departure signals growing burnout among infrastructure builders who feel displaced by rapid AI commoditization and the erosion of craft-oriented work. The move reflects a deeper tension within technical communities: as AI automates and devalues certain skill sets, some veterans are choosing exit over adaptation, raising questions about retention and morale in open-source ecosystems that underpin AI infrastructure.Simon Willison·May 3072
ResearchGenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective TestingResearchers propose GenPT, a psychometric framework that replaces self-report questionnaires with generative projective testing to assess persona-conditioned LLM agents. The work addresses a critical methodological gap: training data contamination and social-desirability bias that plague conventional personality instruments when applied to AI systems. By adapting classical psychology paradigms (TAT, Rorschach, SCT) with procedurally generated stimuli, GenPT offers a more robust measurement approach for evaluating agent behavior and psychological states. This matters for anyone building or auditing persona-driven systems, as it establishes more defensible evaluation standards beyond self-report.arXiv cs.CL·May 3062
ResearchModels & ReleasesMomento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic ConversationsMomento exposes a critical gap in how agentic AI systems handle continuity across user interactions. The benchmark reveals that current agents struggle to distinguish between stale historical context and present-day user state, leading to failures in multi-session task completion where tool use and personalized goals evolve over time. This finding matters because production AI assistants increasingly operate across fragmented sessions, and the field has largely benchmarked single-turn performance. The research signals that persistent memory and temporal reasoning are harder problems than existing evaluations suggest, reshaping how teams should architect agent systems for real-world deployment.arXiv cs.CL·May 3062
ResearchNot All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM DebateResearchers have isolated three distinct mechanisms driving agent convergence in multi-agent LLM debate, revealing that what appears to be productive deliberation often masks social conformity and model instability. The decomposition framework shows 37% of stance shifts stem from self-reflection alone, while strict conformity accounts for 29% of convergence in primary benchmarks. This finding challenges the assumption that debate-based reasoning improves LLM outputs and suggests practitioners must distinguish between genuine persuasion and spurious agreement when deploying multi-agent systems for complex reasoning tasks.arXiv cs.CL·May 3062
Opinion & AnalysisQuoting Daniel JalkutDaniel Jalkut articulates a centrist position on AI adoption that challenges the polarization dominating industry discourse. His framing suggests the productive path forward lies between techno-utopianism and blanket rejection, a stance gaining traction among pragmatist technologists tired of binary framings. This perspective matters because it reflects how informed builders are repositioning themselves as the hype cycle matures and real tradeoffs become visible. For insiders, it signals a potential shift in how the conversation moves from ideological positioning to nuanced capability assessment.Simon Willison·May 3064
ResearchModels & ReleasesCross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMsGoogle's Gemma model family exhibits a counterintuitive safety pattern: the mid-generation Gemma 3 (12B) proves significantly more vulnerable to adversarial attacks than both its predecessor and successor, with attack success rates peaking at 68.7% before dropping to 33.9% in Gemma 4. Using automated red-teaming via quality-diversity evolution, researchers discovered that safety improvements don't scale linearly across model sizes or training iterations. Critically, Gemma 4's defenses generalize beyond the specific attack distributions used in earlier generations, suggesting qualitative shifts in alignment strategy rather than incremental hardening. This non-monotonic pattern has immediate implications for practitioners evaluating model safety claims and for alignment researchers designing robustness benchmarks.arXiv cs.CL·May 3068
ResearchQuality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM SafetyResearchers have developed a quality-diversity evolutionary algorithm that discovers interpretable adversarial attacks against LLMs by maintaining a diverse archive of semantic-level strategies rather than token-level perturbations. Testing across GPT-4o-mini, Claude 3.5 Sonnet, Gemini 2.0 Flash, and Devstral-small-2 revealed distinct vulnerability profiles, with GPT-4o-mini showing susceptibility to hypothetical framing combined with ROT13 encoding. This approach addresses a critical gap in LLM safety testing: manual red-teaming doesn't scale, LLM-as-attacker methods collapse into repetitive patterns, and gradient-based methods produce unintelligible noise. The framework's ability to systematically map behavioral vulnerability landscapes could reshape how organizations prioritize safety interventions.arXiv cs.CL·May 3068
Products & AppsBusiness & Funding‘What a joke’: Github Copilot’s new token-based billing spurs consternation among devsGitHub Copilot's shift to token-based pricing marks a strategic inflection point in how AI coding assistants monetize. Microsoft is moving away from flat-rate subscriptions toward consumption-based models, a pattern that mirrors broader LLM economics where inference costs scale with usage. Developer backlash signals tension between AI vendors seeking margin expansion and users accustomed to predictable pricing. This pricing architecture choice will likely influence how other AI tool makers balance accessibility against profitability as the market matures beyond early adoption.TechCrunch - AI·May 3069
Hardware & InfraProducts & AppsMicrosoft and Nvidia reportedly team up on AI PCs that run actual agents instead of CopilotMicrosoft and Nvidia are jointly architecting a new generation of Windows PCs built around Nvidia's processors and local AI agents powered by the OpenClaw framework, positioning this as a successor to the underperforming Copilot+ initiative. The shift from cloud-dependent Copilot to on-device agentic systems reflects industry recognition that prior consumer AI PC strategies lacked compelling use cases. This partnership signals a strategic pivot toward autonomous task execution on client hardware, with Dell and Microsoft Surface devices expected to debut the approach at Computex and Build conferences, reshaping how OEMs and software vendors compete in the AI-native PC market.The Decoder·May 3080
Hardware & InfraProducts & AppsMeta is reportedly developing an AI pendantMeta is investing in wearable AI hardware, signaling a strategic pivot beyond software toward embodied intelligence devices. The pendant form factor suggests the company is exploring always-on, ambient AI assistants that compete with similar bets from Apple and others in the spatial computing race. This reflects a broader industry shift where major platforms are hedging against mobile saturation by embedding AI into physical form factors, potentially reshaping how users interact with AI systems outside traditional screens.TechCrunch - AI·May 3065
Business & FundingOpinion & AnalysisThe groupthink boom: what 3 top VCs really think about the AI frenzyVenture capital's AI funding frenzy is creating a youth-obsessed talent market where age itself signals capability. The observation that 19-year-old founders attract Series A offers while 22-year-olds compete for seed rounds reveals how compressed timelines and hype-driven valuations are reshaping startup formation. This dynamic signals both opportunity concentration among the youngest builders and potential risk: VCs may be optimizing for narrative momentum over execution track record, creating a cohort of underfunded but older founders and inflating expectations for precocious teams without proven product-market fit.TechCrunch - AI·May 3065
Products & AppsOpinion & AnalysisTerence Tao on How AI Is Changing MathematicsOpenAI and Fields Medalist Terence Tao explored how AI is reshaping mathematical research at an IPAM-partnered forum in March 2026. The conversation centered on AI's role in accelerating experimentation and collaborative problem-solving while maintaining human creativity as the core driver of discovery. Mark Chen, OpenAI's Chief Research Officer, positioned this as part of a broader strategy to scale scientific breakthroughs through AI-assisted tools. The framing signals a shift in how frontier labs view their role: not replacing mathematicians, but expanding the frontier of what's tractable to explore, with implications for how research institutions adopt AI infrastructure.OpenAI (YouTube)·May 3069
Policy & RegulationProducts & AppsAI grifters are creating fake Black people to sell Shein junkSynthetic media generation tools are enabling coordinated fraud at scale, with bad actors deploying AI-generated personas to manipulate social commerce platforms and drive dropshipping sales. The scheme exploits both generative AI capabilities and algorithmic recommendation systems, revealing a critical gap between synthetic content detection and platform enforcement. This represents an emerging class of AI-enabled fraud that combines image generation, identity spoofing, and social engineering, forcing platforms and regulators to reckon with authenticity verification as a core infrastructure problem.The Verge - AI·May 3069
ResearchMaking AI chatbots helpful weakens their ability to simulate human behavior, large-scale study findsA large-scale empirical study tracking 208,000 participants across 26 million responses reveals a fundamental tension in language model development: the alignment techniques that make models safer and more helpful systematically degrade their capacity to predict human behavior patterns. The degradation compounds across model generations, suggesting that helpfulness training and behavioral fidelity operate as opposing objectives. Even demographic persona injection, a common industry workaround, yields negligible gains for individual-level prediction accuracy. This finding challenges assumptions underlying human-AI interaction research and raises questions about whether current alignment approaches inadvertently push models away from human-like reasoning.The Decoder·May 3080
Opinion & AnalysisResearchTerence Tao argues AI could bring division of labor to math for the first time in historyTerence Tao's vision of 'industrial mathematics' marks a conceptual shift in how mathematical research could be organized at scale. Rather than individual researchers mastering every phase of problem-solving, AI-augmented teams could specialize in discrete tasks, with humans retaining authority over high-level intuition and conjecture. This mirrors broader labor restructuring across knowledge work, but carries particular weight in mathematics, where the lone-genius model has dominated for centuries. The implication for AI infrastructure is substantial: demand for systems that can handle verification, computation, and proof-checking at scale, while remaining transparent enough for human mathematicians to trust and guide.The Decoder·May 3073
Products & AppsPolicy & RegulationAttackers abuse shared ChatGPT and Claude chats to spread malwareThreat actors are weaponizing the share-link feature in ChatGPT and Claude to distribute malware payloads disguised as legitimate error messages or setup instructions. Because these conversations live on Anthropic and OpenAI's trusted domains, they bypass traditional email and web filters, creating a new attack surface that exploits user trust in first-party infrastructure. This signals a shift in how LLM platforms themselves become distribution channels for social engineering, forcing both companies to rethink access controls and content moderation on shared artifacts.The Decoder·May 3073
Products & AppsTools & CodeOpenAI's Codex can now operate your Windows PC autonomously, hunting bugs and testing apps on its ownOpenAI has extended Codex capabilities to Windows 11 with autonomous computer control, enabling the model to independently execute software testing, bug detection, and application validation without human intervention. The integration includes remote task initiation and monitoring via ChatGPT mobile, marking a significant expansion of AI agent autonomy beyond code generation into full system operation. This development signals movement toward practical agentic AI in enterprise workflows, though it also raises questions about security, oversight, and the operational risks of unsupervised model access to production systems.The Decoder·May 3085
Products & AppsBusiness & FundingSalesforce claims AI agents cut a 231-day migration to 13 days with fewer incidentsSalesforce's migration of its development infrastructure to Anthropic's Claude Code reportedly compressed a 231-day project into 13 days, with developers shipping 79 percent more pull requests and incident rates dropping five percent. The case crystallizes a fault line in engineering culture: whether AI agents represent genuine productivity transformation or a new vector for technical debt accumulation. The unverified metrics matter less than what they signal about enterprise adoption velocity and the stakes vendors see in the agentic coding shift.The Decoder·May 3073
Products & AppsBusiness & FundingMeta's leaked memo reveals AI pendant, supersensing glasses, and enterprise wearables strategyMeta is pivoting from software-first AI strategy to hardware-centric deployment, signaling a major shift in how the company plans to monetize its AI investments. The leaked roadmap outlines AI-enabled wearables including a pendant and advanced glasses with sensing capabilities, alongside enterprise hardware products. This move reflects broader industry recognition that AI's commercial value increasingly depends on embodied form factors and real-world data collection rather than model scale alone. For investors and technologists, the strategy suggests Meta believes consumer and enterprise wearables represent the next frontier for AI adoption, potentially reshaping competition in spatial computing and edge AI.The Decoder·May 3073
ResearchOpinion & AnalysisCoders are refusing to work without AI , and that could come back to bite themDeveloper reliance on AI coding assistants is reshaping workforce expectations, but emerging research suggests speed gains may mask quality degradation. This tension between productivity metrics and code robustness creates a hidden technical debt problem: teams optimizing for velocity risk shipping fragile systems that compound maintenance costs later. The trend signals a critical inflection point where AI adoption outpaces organizational maturity in evaluating actual output quality, forcing engineering leaders to recalibrate how they measure AI-assisted development success.TechCrunch - AI·May 2969
Policy & RegulationBusiness & FundingAmazon Is Making an AI-Animated ‘Good Advice Cupcake’ TV Show. Its Original Creator Is FuriousAmazon's use of AI animation to produce a licensed TV series without the original creator's consent crystallizes a growing tension in media production: as generative tools lower production costs, IP holders and studios can now bypass creator involvement entirely. This case sits at the intersection of copyright enforcement and AI-enabled workflow disruption, raising questions about consent frameworks when AI becomes the production layer. For the industry, it signals that licensing agreements written before AI animation matured may lack sufficient protections, forcing creators and studios to renegotiate terms around synthetic media generation.WIRED - AI·May 2969
Products & AppsResearchHands-On With Gemini Spark: I Gave It Access to My Life and It Friend-Zoned My BoyfriendGoogle's Gemini Spark agent represents a shift toward AI systems that operate autonomously across personal data streams, yet this hands-on test reveals a critical gap in contextual reasoning. By accessing emails, documents, and calendars to execute a real-world task (party planning), the agent demonstrated capability limitations in understanding relational hierarchies and social context, despite having raw data access. The failure to identify a user's primary relationship exposes how current agentic systems struggle with implicit human priorities, raising questions about whether data breadth alone translates to meaningful personalization in high-stakes domains.WIRED - AI·May 2965
Products & AppsTools & CodeWindows Computer Use and mobile access for CodexOpenAI has expanded Codex's computer-use capabilities to Windows, enabling the agent to operate desktop applications autonomously while users are away, and introduced remote control via ChatGPT mobile. This represents a meaningful step toward practical agent deployment beyond chat interfaces, shifting the value proposition from conversational assistance to delegated task execution across heterogeneous software environments. The mobile-to-desktop bridge signals OpenAI's strategy to embed agentic workflows into everyday device ecosystems, raising questions about how enterprises will govern autonomous desktop access and what new security and compliance challenges emerge when LLM agents interact with legacy Windows applications at scale.OpenAI (YouTube)·May 2969
Models & ReleasesProducts & AppsOpenAI gives GPT-5.5 Instant a readability upgrade while phasing out two older modelsOpenAI is refining GPT-5.5 Instant's output naturalness while consolidating its model lineup, retiring o3 and GPT-4.5 by August 2026. The shift also eliminates Canvas, moving writing and coding workflows directly into chat. This reflects OpenAI's strategy to streamline its product surface and push users toward its latest generation, signaling confidence in 5.5's capabilities across diverse tasks while reducing support overhead for aging models.The Decoder·May 2968
Business & FundingOpinion & AnalysisWhat happens when companies become too AI-pilled?Aaron Levie's critique of 'AI psychosis' surfaces a structural problem in enterprise automation: executives deploying AI agents to eliminate roles often lack operational visibility into what those roles entail. ClickUp's 22% workforce reduction exemplifies this pattern, with 2026 tech layoffs already tracking toward 2025 totals. The tension reveals a gap between AI capability and organizational wisdom, forcing insiders to reckon with whether efficiency gains justify the collateral damage of misaligned automation decisions.TechCrunch - AI·May 2969
Products & AppsBusiness & FundingGoogle fixes several bugs in Gemini usage limits that burned through quotas too fastGoogle has patched critical quota-management flaws in Gemini that allowed single video generations to exhaust entire user allowances. The fixes include doubling video generation limits for Ultra subscribers, eliminating charges for failed requests, and forthcoming usage transparency improvements. This incident underscores the operational friction emerging as multimodal AI tools scale, where billing systems and rate-limiting infrastructure lag behind capability deployment. For power users and enterprise adopters, quota predictability directly impacts adoption velocity and willingness to commit to paid tiers.The Decoder·May 2968