CodeRabbit 官方最新动态:The Merge: Your tests pass. Can your agent finish the job?
来源:CodeRabbit 官方动态 | 发布日期:2026-09-22
核心更新概览
In CodeRabbit’s latest episode of The Merge, I sat down with Francesco Bonacci, CEO and co-founder of Cua AI, which builds tools and environments that let AI agents operate computers. We talked about
详细内容记录
In CodeRabbit’s latest episode of The Merge, I sat down with Francesco Bonacci, CEO and co-founder of Cua AI, which builds tools and environments that let AI agents operate computers. We talked about how his team tests those tools and checks whether agents can use them to complete an assignment. A test can establish that an agent’s click reaches the intended button. It cannot, by itself, establish that the agent chose the right button or finished the job. Cua’s engineering team tests whether computer controls work, whether agents complete their assignments, and whether the software grading those assignments scores them correctly. Bonacci described the challenge of catching regressions in computer control across different operating environments. The release pace adds pressure. As he put it, “we pushed out 21 minor releases in one month and a half.” explains that its test applications independently observe whether an action produces the expected change. The test must confirm what happened in the application, even when the tool reports success. Check the work the agent leaves behind To check whether an agent completed an assignment, the team uses Cua-Bench. In Cua’s separate AI Engineer World’s Fair presentation , CTO Dillon DuPont explains that each task has a known starting state, a reference solution, and an evaluator. The evaluator is software that examines files or application state to determine whether the agent succeeded.
更多技术细节可访问官方原文:https://www.coderabbit.ai/blog/the-merge-cua-can-your-agent-finish-the-job。