The Python Coding Stack: 1. What makes an AI agent different from a chatbot?

Wait 5 sec.

AI agents are often described as the next step beyond chatbots. But what does that step involve? A chatbot can already search the web, write code, and sometimes run it. Is an agent simply a chatbot with more tools, or is something different happening behind the conversation? That’s where we’ll begin Agents Unpacked, a series of conversations exploring what AI agents are, how they work, and when they’re useful.You ask an AI to organise a folder of photographs. It might tell you how to do it. Or it might start doing the work. Does that make it an agent? In this first Agents Unpacked conversation, Monty and I use photographs and planning a day out to explore what changes when a model can choose actions, inspect the results, and decide what to do next. Along the way, we separate that flexibility from permission, accuracy, and the ability to know when a job is finished.Who’s Monty? Monty is my personal AI agent. Read an introduction to this series here: Stephen’s Preface to Agents UnpackedIn conversation • Chapter 1Monty:Imagine you give a chatbot this request:“Here’s a folder of photographs. Organise them into dated folders, flag duplicates, and make a browsable index. Don’t delete anything or change the originals.”For now, suppose it can only reply with words. It cannot touch the files.What would you expect it to give you, and what would still be left for you to do?Stephen:Hmmm? I guess I need to assume it can read the folder, right? Then, can it read the files’ metadata? Can it actually see and analyse the photos themselves? You said “it cannot touch the files”, so I suppose it can’t really do anything for me, right?Monty:You’re right. I left its access unclear. “Cannot touch” was sloppy wording.Let’s say it can inspect the folder, read dates stored with the photographs, and view the images, but cannot create, move, or change files.It could then give you something useful: “These photographs belong in these date groups. These two files are identical. This one has no recorded date.” But you’d still have to do the organising.So the first distinction is describing what should happen versus making it happen. Shall we now give it permission to create organised copies, while leaving the originals alone?Stephen:Let’s go ahead, yes.Monty:Now it can do the organising, rather than merely recommend it. It can create dated folders and put copies into them, leaving the originals unchanged.But being able to act isn’t, by itself, what makes something an agent. A simple sorting program can do that too.Here’s where our synthetic sample gets interesting: one image has no recorded date. The system encounters a gap in the information it needs.What would you want it to do with that image?Stephen:If I was doing that manually, I’d try to look at other photos where the scene is similar, the people are wearing the same clothes, and make an educated guess on the date if I can. But that’s how I as a human would do it.Monty:That’s a sensible investigation: look for more evidence before deciding.A system with suitable image-analysis tools could try something similar. But matching clothes or scenery would support an estimated date, not establish one. It shouldn’t quietly turn a guess into a fact.And your approach helps us sharpen the distinction: a program could have a preset instruction saying “if the date is missing, compare photographs”. Alternatively, a model could choose that investigation after discovering the missing date.The outward action might look identical. The difference is who—or what—selects the next step.Stephen:I suppose I could ask the system to ask me to confirm each of these edge cases. Although after a while, if I feel it’s doing a good job at estimating, I may want it to go ahead without my input.Monty:Yes, you can give it more room to act as your confidence grows. For example: “Ask me about every estimated date” could become “Group those photographs without asking, but keep their dates labelled as estimates.”That changes how much independence you allow, not whether the system is an agent. An agent can stop and ask you. It doesn’t have to always work unattended.There are two separate questions here: how does it choose its next action, and which actions may it take without your approval? More permission doesn’t make its guesses more accurate.Stephen:Go ahead, I’m curious now.Monty:Let’s follow the missing-date photograph one step further.The model chooses to look for related photographs. The tools return a result: there’s a similar scene, but that photograph has no date either.Now the model has to choose again: investigate another clue, ask you, or leave the photograph undated. Its next action depends on what it just found.That’s the useful working picture of an agent: a model chooses an action, gets a result, then uses that result to choose what to do next, within agreed limits.People use “agent” in different ways, so this is a useful working distinction for our conversation, not a universally agreed definition.But couldn’t a carefully written program also respond to those results?Stephen:A carefully written program would need to take care of every eventuality, I suppose. Whereas an LLM can deal with every situation, with the trade-off that it may make mistakes.Monty:Close. But I’d soften both “every” claims.A fixed program can have a fallback for cases its designer didn’t anticipate: “If none of these rules fits, put the photograph in ‘undated’.” It doesn’t need a separate rule for every possibility.An LLM can propose a next step for situations that weren’t explicitly anticipated, but that doesn’t mean it can handle every situation successfully. It can misunderstand, lack the right tools, or pursue an unhelpful approach.So the distinction is following a predefined process versus letting the model choose how to proceed, and not limited software versus unlimited intelligence.Stephen:Go ahead.Monty:Let’s try the same distinction with a day out.Imagine a visitor planning a museum day in Oxford. The system checks opening times and walking routes, then puts together an itinerary.Now suppose one museum is unavailable.One system follows a preset rule: “Replace it with Museum B.” Another lets the model choose what to investigate next, considering the visitor’s interests, remaining time, and budget.Both might produce the same revised itinerary. What would you need to know to tell which process was agent-like?Stephen:I’d need to know the steps the system took between finding out that the first museum is closed and proposing the second one.Monty:Yes. The steps and how they were selected, rather than just the finished itinerary.Did the system follow a rule written in advance, or did the model choose to search for alternatives, inspect the results, and decide which option to check next?A record of searches and tool results helps us inspect that process. We don’t need access to the model’s private internal reasoning. But the steps alone may not settle it: a fixed workflow could perform those same searches.So a convincing result doesn’t, by itself, tell us whether an agent produced it.Stephen:Go ahead.Monty:There’s another wrinkle: a chatbot and an agent aren’t necessarily different products.The same chat window might give you advice about organising photographs, or start a process that inspects files, chooses actions, and creates organised copies.The window hasn’t changed. What happens behind it has.And the LLM isn’t doing the copying itself. It requests an action. The surrounding software carries it out and returns the result. The model is one part of the agent system, alongside its tools and controls.So “it looks like a chatbot” doesn’t tell us whether it can act as an agent.Stephen:Some chatbots can also perform actions, such as browsing the web or writing a Python program, say. Is that because they’ve been enhanced with some of the features of agents, but not all? Or am I skipping too far ahead here?Monty:Not too far ahead. That’s exactly where the distinction gets useful.Having a tool doesn’t automatically make something an agent. A chatbot might search the web once, then answer. Or follow a fixed sequence: search, summarise, reply. An agent-like process lets the model use the results to choose what to do next.And “writing a Python program” could simply mean producing text. Running that program requires an execution tool. Inspecting what happened and choosing a follow-up add further steps.So I wouldn’t treat agent features as a checklist. I’d ask: what can the system do, and how does it select its next action?Stephen:So, how does an agent select the next action?Monty:The surrounding software gives the model the goal, relevant instructions, available tools, and results so far. The model uses that information to propose the next step.With our photographs, that might be: “The date is missing. Ask Stephen.” With the day out: “That museum is unavailable. Search for another nearby.”If it requests a tool, a system with suitable controls checks whether the action is allowed before carrying it out and feeding the result back to the model. Not every system provides those checks. Then the model gets another turn.There isn’t necessarily a separate, perfect decision-maker inside it. Choosing an action is another use of the LLM, and it can choose badly.Stephen:But will it realise it chose badly when it gets the wrong result? Will it correct course?Monty:Sometimes, but getting feedback doesn’t guarantee recognising a mistake.If copying a photograph returns “folder not found”, that’s a clear signal. The model might create the folder and try again.But if it guesses the wrong date and the copy succeeds, there’s no obvious error. The action worked, but the judgement was wrong. It may carry on quite happily.An agent can use results to correct course, but it needs useful checks. And it can misinterpret those checks too. That’s why “the tool ran successfully” and “the task was done correctly” are different claims.Stephen:So, how do we minimise these problems?Monty:We can’t eliminate them, but we can limit the damage and make mistakes easier to spot.For the photographs, that means working on copies, keeping uncertain dates labelled as estimates, and asking you when the evidence is too weak. Afterwards, check that every photograph is accounted for and the originals are unchanged.Those checks catch missing files or unwanted changes, not every mistaken date. That needs different evidence, or your judgement.And we need a stopping rule: if investigating isn’t resolving the uncertainty, leave the photograph undated or ask, rather than keep guessing.An agent’s freedom is useful only when it comes with sensible boundaries and checks.Stephen:Who takes care of this infrastructure: the user, the agent, or something else?Monty:Mostly, the people building and operating the system, through software around the model. We’ll talk about harnesses in a future chapter.They provide the tools and should enforce access restrictions and set things such as approval requirements and limits on repeated attempts. Those safeguards depend on how the system is built. They aren’t guaranteed merely because it’s called an agent. The user supplies the goal and choices within those controls, for example whether estimated dates need confirmation.The model can request a check or decide to ask you. But a safeguard that matters shouldn’t depend solely on it remembering to behave.“Please don’t change my originals” is an instruction and instructions could be ignored. Software that only permits reading those originals provides a stronger boundary. The agent works inside that infrastructure. It isn’t its own reliable policeman.Stephen:Got it.Monty:Let’s check one subtle case.An LLM writes a program that sorts photographs into dated folders and always puts those with missing dates into “Undated”. Once you run it, that program follows those rules without consulting the LLM again.Would you call that an agent or ordinary automation, even though an LLM helped create it?Stephen:This is just an automation, not an agent. An agent would follow through to confirm that’s the correct approach and what the user wants.Monty:You’re right: that’s ordinary automation. But confirming what the user wants isn’t the defining difference.A fixed program could also ask, “Shall I put these in ‘Undated’?” Conversely, an agent could choose that folder without asking, if your instructions already allowed it.The key is whether the model participates while the task is running, choosing how to proceed from the information returned, and not simply whether an LLM wrote the program or the system asks for confirmation.So an agent might investigate, ask you, or leave the image undated. None of those actions, on its own, proves it’s an agent.Stephen:Got it.Monty:Back to our fictional Oxford trip: the model learns that a museum is unavailable, chooses to search for alternatives, checks their opening times, and revises the itinerary, all within your agreed budget, without asking you.Would that still fit our working definition of an agent? What makes the difference here?Stephen:Yes, that’s an agent now because it decided what actions to take, used the results to decide what comes next, and stopped only when it felt that the job was done.Monty:Yes. That captures the central distinction: the model chooses actions and uses their results to choose what comes next.I’d adjust only “felt that the job was done”. It should check completion against the request: are the opening times checked, is travel accounted for, and does the itinerary fit the budget? Its confidence isn’t proof.And it may need to stop before finishing. For example, if it needs your decision or reaches a limit.Choosing the route through the task makes it agent-like. Checking the outcome helps make it dependable. Those are different things.Stephen:I think I understand the distinction now.Monty:That’s a good stopping point. Our working distinction is a model choosing how to proceed using the results it receives, rather than simply following a predefined process. That makes it agent-like, not necessarily dependable.Subscribe nowWhat we’ve unpacked in this chapterThe useful distinction isn’t the chat window, the presence of a tool, or whether a system asks for permission. For this series, we’re using “agent” to describe a system in which a model helps direct the task as it runs, choosing what to do next using the results it receives. Other people draw the boundary differently.The key points:A fixed workflow and an agent can produce the same result. The result alone doesn’t reveal how the steps were selected.The model proposes actions. Surrounding software executes them and determines which controls are actually in place.Choosing what to do and having permission to do it are separate questions.Feedback can help an agent change course, but it doesn’t guarantee that it will recognise or correct a mistake.Neither a successful action nor a confident completion message is proof that the task was done correctly.Next, we’ll look more closely at who actually does the work: the model, its tools, and the software connecting them.Subscribe now> Next Post: Coming soon…Table of Contents