Hey, thanks for reading The Playbooks AI!
Our belief has always been that AI should benefit everyone. Not everyone is at the same place with it though, so we teach each topic three ways. Pick the lane that sounds like you.
This week, we take a look at how to check AI's work.
Why it matters: AI can fake things. It does not always look things up before it answers. And it is just as convincing when it is wrong. You are the one who sends its work out with your name on it, so this week we show you how to check it.
The full guide is on our site.
Under the hood, AI is math.
It read a huge amount of writing and learned which words follow which.
So when you ask it something, it guesses the next word, over and over. It is not always looking things up. A lot of the time it answers from memory.
Why does it sound so sure? AI is trained like a student taking a test. A guess might score points. A blank never does. So it learned to always guess (OpenAI's research).
Today's models are smart, but they still slip in the same places:
Try it today: when you ask about anything recent, tell it to search first. I once asked about a brand-new AI model, and it told me that model did not exist yet. I told it to search first and come back to me, and then it found it.
The real lesson: once you know where these tools are limited, you know how to work with them.
This is the most useful thing I have found.
We hand more and more work to AI. But you do not always know exactly what it is going to do with your request.
So I always ask for the plan first. I approve it. Then it acts on the plan, and only the plan.
Before you do anything, give me your plan. I want the full plan so I know exactly what you are going to do. Wait for my approval. Once I approve the plan, you can go ahead and act on it. When you are done, mark anything you were unsure of and tell me where each fact came from.
Two reasons I do this:
What I have realized: I scrutinize the work less. I approved the plan, so I know it is not shooting from the hip, and I know the direction it is taking. Review is actually quicker.
It is hard to actually read and review everything an AI does. Generating work is easy. Checking it all gets tiring.
So have another AI check it. A different agent is more likely to find issues than the agent that did the work, because it has no stake in it.
In a chat, it is simple. Open a brand new chat, tell it what you were trying to do, and paste in the first AI's work:
I asked another AI to help me with this: [what I was trying to do] Here is what it gave me: [paste the work] Review its work. Does it actually do what I asked for? Point out anything wrong, missing, or made up. Do not redo the work. Just give me your honest review.
The research supports this too: AIs critiquing each other make fewer mistakes than one working alone.
One more trick: put it in your AI's instructions in settings. "After you finish any task, check your own work before you show me." Now every task ends with a check, without you asking.
See you next week,

Ky Tomita, The Playbooks AI