← All articles

ARTICLE · TODD KELSEY

Agent Wars VI: Ping Pong — A Panel of Expertise with Two AIs

I once made a small game called Emoji Pong. Two paddles, a snack for a ball, one player facing a friend or a computer. The useful part of ping pong is that the ball comes back. A second AI can do that for an idea.

For Agent Wars V, I drafted an argument about making money with AI agents and sent the central claim to Gemini. It found the rejected cases I had undercounted. Its own break-even explanation then conflated variable costs with fully burdened costs. I returned the correction. On the next pass it produced useful sensitivity arithmetic, along with unsupported market rates and arbitrary pilot thresholds. I returned those too. By the third pass the thesis was narrower and more honest.

This is an AI panel in a modest, practical sense: different systems, prompted to play different roles, with a person setting the question and choosing what survives. It is not a panel of credentialed human experts, and agreement between models is not evidence that a claim is true.

Give each pass a job

  1. Builder: Draft a concrete claim with sources, assumptions, and a worked example.
  2. Adversary: Ask another model for the strongest counterexample, omitted costs, faulty arithmetic, and sentences that sound more certain than the evidence.
  3. Referee: Check the objection against primary sources and a calculator. Reject unsupported objections too.
  4. Rebuilder: Revise the piece and ask the second model to attack the revision once more.
  5. Human editor: Decide what matters, what can be published, and what still needs a real-world test.

One AI can use another when tools or authorized workflows permit it; otherwise a human can carry the draft and critique between two chats. In either arrangement, pass only the material needed for review, retain the prompts and outputs, and label which claims were checked independently. A model can quote another model's error back to it with great confidence. More voices do not automatically mean more independent evidence.

The exercise works best with a visible ledger: claim, supporting source, challenge, response, unresolved question. For the money article, the ledger exposed an important distinction: an illustrative $940 monthly remainder was not proof of an attractive business. It also let me show where Gemini itself had made a numerical mistake. A review that cannot be corrected is theater.

A small class could try the method with a game design. Ask one AI to build a simple browser game; ask another to test the controls on a phone, look for inaccessible text, and spot ways it could be more fun. Then play it yourself. Emoji Pong is the Easter egg here, and a reminder that the user gets the final turn.

The aim is neither automatic consensus nor endless debate. Stop when the claim is clear enough to test with the world, or when the next exchange is merely rephrasing the last. Good ping pong ends with someone stepping away from the table and doing the work.

Read the companion economics article, Agent Wars V →