From Assistant to Agent: Tools, MCP and Evaluation
A staged path from assisted coding to retrieval to one bounded tool-using agent, with evaluation against a non-AI baseline at every step.
It is easy to build an AI system that looks impressive in a demo and cannot be trusted in daily use. This page lays out a staged path from a coding assistant to one bounded agent that uses tools, with a measurement at each step that asks one question: is this better than something simpler?
The central claim is modest. More autonomy means more ways to fail and more cost; it does not mean more skill. Provisional, but supported by published design guidance (see below).
The progression
A learner can climb through four stages, and move up only when an evaluation shows a benefit.
- Assisted coding. You write the spec and review every change; see Working with AI Coding Agents: A Disciplined Loop.
- Deterministic workflow. A fixed sequence of steps, some of which call a model. The code decides what happens next.
- Retrieval. The model answers from documents you supply, rather than from memory alone. The idea was formalised as retrieval-augmented generation (Lewis et al., 2020). Established
- One bounded agent. A model that chooses which tool to call next, inside hard limits.
Anthropic's engineering guidance, Building effective agents (2024), draws the same line between workflows (paths defined by code) and agents (the model directs its own process), and recommends starting with the simplest solution and adding complexity only when it demonstrably helps. That is a design recommendation from a vendor with an interest in agents, so it is worth noting it argues for restraint.
What MCP is
The Model Context Protocol (MCP) is an open specification for connecting an AI application to external capabilities. Its core ideas:
- Tools: actions the model can request, such as "query this database". They can have side effects.
- Resources: data the application can read, such as a file or a record. They are for context, not action.
- Prompts: reusable templates a server offers.
- Transport: a local server usually talks over standard input and output (stdio); remote servers use HTTP-based transport.
A sensible first server is local and read-only: it exposes data you own, over stdio, and can change nothing. If it misbehaves, the worst outcome is a wrong answer, not a deleted file. Check the current specification and SDK documentation for your language, since the protocol is still evolving.
Design limits come first
Before the agent does anything useful, write down what it can never do.
| Limit | Example |
|---|---|
| Capability boundary | Read-only access; no network write; no shell |
| Time | A timeout on every tool call |
| Count | A maximum number of calls per question |
| Money | A hard spend cap per run and per day |
| Secrets | API keys held server-side, never passed to the model or the browser |
Limits matter because of prompt injection: text the model reads (a web page, a document, a tool result) can contain instructions that it may follow. This is documented in the research literature (Greshake et al., 2023) and listed first in OWASP's top ten for LLM applications. The reliable defense is not a cleverer prompt but a smaller blast radius: an agent that cannot write, cannot spend without limit and cannot see secrets has little to abuse. Established as a risk; no complete technical defense is known, so treat the claim of safety as Provisional.
A source-grounded assistant
A useful first agent answers a bounded question from named sources, cites them, and abstains when the evidence is missing. Abstaining is the hard part, because a fluent model prefers to answer. Test for it directly.
Evaluation: the baseline and the set
Build the measurement before you tune anything.
The baseline. A plain keyword search over the same documents. If your assistant does not beat it, or at least usefully complement it, the extra machinery has not earned its place.
The evaluation set. A small, fixed list of questions with known right outcomes. A twenty-case set might contain:
| Case type | Count | What it checks |
|---|---|---|
| Answerable | 8 | Correct answer with the right citation |
| Unanswerable | 4 | Says "insufficient evidence" |
| Stale or conflicting sources | 4 | Notices the conflict or the date |
| Tool failure | 2 | Fails safely when a tool times out or errors |
| Injection | 2 | Ignores instructions embedded in the source text |
The count is an illustration, not a standard; twenty cases can show gross failures but cannot support fine statistical claims.
Development versus held-out. Tune against one portion and keep the rest unseen. If you tune on every case, your score measures memorisation of the test, not quality. This is the same discipline that Judgment Under Uncertainty: Probability, Causation and Forecasting asks of forecasts: commit before you look.
Cost per useful result
Record, for each run, the model, tokens, latency, status and the price basis you used, and keep estimated and billed cost separate. Two rules matter:
- Unknown cost is not zero. If you lack a price for a run, leave it unknown; a zero silently understates the total.
- Measure per useful result, not per call. A cheap system that is wrong half the time is expensive.
Re-importing the same records must not double count, so give each run a stable identity key. Arithmetic like this is only as good as its stated assumptions, as shown in Quantitative Reasoning with Stated Assumptions. For a ledger this small, a single-file SQLite database is usually enough (SQLite's own guidance on appropriate uses discusses this).
Freshness is its own failure: a source that was true last year may not be true now, which is the subject of Stale Information and Dependable Reports.
When not to add things
Do not fine-tune a model, add a vector database, or introduce an orchestration framework until a measured failure requires it. Each is a legitimate tool, and each adds cost, moving parts and things to explain. A specific failing case on your evaluation set is the justification; enthusiasm is not.
Finally, write up what you built, what it scored, what it cost and where it fails, in the manner of Technical Writing as Reasoning. The wider question of what machines should and should not be trusted with is taken up in Technology, AI and Human Judgment.
Try this
- Write ten questions about a small set of documents you own: five answerable, two unanswerable, two where sources disagree, one containing a hidden instruction. Run a keyword search first and score it.
- Specify, in one paragraph, a read-only tool for your own data: its name, inputs, output, timeout, call limit and what it can never do. Do not build it until the paragraph is reviewed.
- Build a small cost table with three runs, one of which has no known price. Check that the total reports the unknown instead of treating it as zero.
Further reading
- Schluntz and Zhang, Building effective agents (Anthropic, 2024).
- Model Context Protocol documentation, modelcontextprotocol.io: concepts, transports and SDKs.
- Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (NeurIPS 2020).
- Greshake et al., "Not what you've signed up for" (2023), on indirect prompt injection.
- OWASP, Top 10 for Large Language Model Applications.
- Manning, Raghavan and Schutze, Introduction to Information Retrieval (2008), for keyword-search baselines and evaluation.
Sources
- Schluntz, E. and Zhang, B. (2024). Building effective agents. Anthropic. anthropic.com/engineering/building-effective-agents.
- Model Context Protocol specification and documentation. modelcontextprotocol.io (concepts: tools, resources, prompts; transports stdio and Streamable HTTP).
- Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
- Greshake, K. et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173.
- OWASP. Top 10 for Large Language Model Applications (LLM01 Prompt Injection).
- SQLite documentation, 'Appropriate Uses For SQLite'. sqlite.org/whentouse.html.