Install
The New Stack is a media platform for the people who build and manage software the world relies on. We provide context and explanation of at-scale technologies to advance knowledge and create conversations through our coverage of modern architectures, components of the software development life cycle, and operations to
- 154articles · 30d
- 2+ hour agolatest article
- Aug 14, 2026earliest in window
- 96%with images
- 87avg words
- Science & Technology 142
- Software Dev. 102
- Computers & Electronics 96
- News 39
- Software 21
- Science & Nature 13
- Economy, Business & Finance 9
- Finance & Business 9
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
It passed CI. It passed your evals. The customer still got the wrong answer.
3+ hour, 24+ min ago (771+ words) Your AI agent returned a 200, passed its faithfulness check, and still answered the wrong question. The evidence that explains why lives in the trace....
The AI-native SDLC won't be one process
1+ day, 3+ hour ago (92+ words) Anthropic says code is no longer the bottleneck. It's right -- but the process that catches your agent's mistakes can't be one size for every change....
AWS open-sources Pizza Bot: email-style inbox for background AI agents
2+ day, 18+ hour ago (280+ words) Two thousand people inside Amazon used early versions - before Pizza Bot was rebuilt as a standalone community project....
Stop AI code sprawl before it destroys your software design
3+ day, 4+ hour ago (450+ words) Prevent AI code sprawl and Comprehension Debt. Use Python tools like pytest-archon to enforce Executable Architecture in your CI/CD....
Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.
3+ day, 21+ hour ago (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....
Building trust in agentic RAG starts with evidence
1+ week, 1+ day ago (500+ words) Agentic RAG requires clear evidence. Discover how tracking retrieval decisions, metadata, and citations builds trust in AI agent outputs....
AI agent evaluations are part of the product
1+ week, 2+ day ago (886+ words) Move beyond simple AI demos. Build repeatable evaluation systems, test execution paths, and enforce strict release gates for AI agents....
It cost $33 to build a virtual Union Square. Here's what the agents got wrong.
1+ week, 2+ day ago (152+ words) AI coding agents built a 3D browser replica of San Francisco's Union Square in two hours, then used Playwright screenshots to catch visual errors no test could....
Want to scale AI agents without breaking anything? Retrieval engineering is the answer.
1+ week, 3+ day ago (23+ words) On September 24, GigaOm’s Whit Walters and Vespa.ai’s Bonnie Chase will explain what breaks when hundreds of AI agents hit retrieval at once....