<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Jonathan Mukhobe | Blog</title>
        <link>https://www.jonathanmuk.com/insights</link>
        <description>Notes written while building AI systems: the Model Context Protocol, retrieval and vector memory, load balancing and rate limiting.</description>
        <lastBuildDate>Mon, 20 Jul 2026 00:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <language>en</language>
        <image>
            <title>Jonathan Mukhobe | Blog</title>
            <url>https://www.jonathanmuk.com/og/pages/insights.jpg</url>
            <link>https://www.jonathanmuk.com/insights</link>
        </image>
        <copyright>© 2026 Jonathan Mukhobe</copyright>
        <atom:link href="https://www.jonathanmuk.com/rss.xml" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[Model Context Protocol: The USB-C Port for AI Applications]]></title>
            <link>https://www.jonathanmuk.com/insights/model-context-protocol</link>
            <guid isPermaLink="false">https://www.jonathanmuk.com/insights/model-context-protocol</guid>
            <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[A complete guide to MCP: servers, clients and hosts, the three primitives, JSON-RPC message shapes, the lifecycle handshake, transports and versioning.]]></description>
            <content:encoded><![CDATA[<h2 id="model-context-protocol-mcp"><a href="#model-context-protocol-mcp">Model Context Protocol (MCP)</a></h2>
<p>This is an open-source standard for connecting A.I applications to external systems. Using MCP, A.I applications like Claude or ChatGPT can connect to data sources (e.g. local files, databases), tools (e.g. search engines, calculators) and workflows (e.g. specialized prompts), enabling them to access key information and perform tasks.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p>Think of MCP like a USB-C port for AI applications. Just as USB-C provides a standardized way to connect electronic devices, MCP provides a standardized way to connect AI applications to external systems.</p></div></div>
<h2 id="the-problem-that-mcp-exists-to-solve"><a href="#the-problem-that-mcp-exists-to-solve">The Problem That MCP Exists to Solve</a></h2>
<p>Before MCP existed, here is what happened every time you wanted to give an AI agent access to an external tool or system. Say you are building a LangGraph agent and you want it to be able to:</p>
<ul>
<li>Read files from Google Drive</li>
<li>Create tasks in Notion</li>
<li>Look up employee records in HiBob</li>
<li>Send Slack messages</li>
</ul>
<p>For each one of those, you had to write a completely custom function, handle authentication your own way, manage errors your own way, format the data your own way, and wire it all up manually. If you later switched from LangGraph to a different agent framework, you would have to rewrite everything. If someone else built a better Google Drive connector, you could not easily plug it in; theirs was written for their system, not yours.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p>This is what Anthropic called the <strong>M×N problem</strong>: M different AI models or apps, each needing to connect to N different tools. Without a standard, you need M×N custom integrations. That does not scale.</p></div></div>
<h2 id="why-does-mcp-matter"><a href="#why-does-mcp-matter">Why Does MCP Matter?</a></h2>
<p>Depending on where you sit in the ecosystem, MCP can have a range of benefits.</p>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 13rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Developers</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>MCP reduces development time and complexity when building, or integrating with, an AI application or agent.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">AI Applications or Agents</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>MCP provides access to an ecosystem of data sources, tools and apps which will enhance capabilities and improve the end-user experience.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">End-Users</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>MCP results in more capable AI applications or agents which can access your data and take actions on your behalf when necessary.</p></div></li></ul>
<h2 id="the-three-roles-in-mcp"><a href="#the-three-roles-in-mcp">The Three Roles in MCP</a></h2>
<p>Every MCP setup has exactly three roles. Understanding these is the foundation of everything else.</p>
<h3 id="mcp-server"><a href="#mcp-server">MCP Server</a></h3>
<p>This is what exposes tools or data to the AI. A program that provides context to MCP clients. It is a program that says &quot;here are the things I can do.&quot; For example, a Notion MCP server says: &quot;I can read pages, create pages, update databases, search content.&quot; A HiBob MCP server would say: &quot;I can get employee records, check leave balances, update profiles.&quot; The server does not know anything about AI; it just knows how to expose its capabilities in the MCP standard format.</p>
<h3 id="mcp-client"><a href="#mcp-client">MCP Client</a></h3>
<p>This is the component inside your AI app that talks to the server. It is the bridge. It asks the server &quot;what can you do?&quot; and then makes requests on behalf of the AI. A component that maintains a connection to an MCP server and obtains context from an MCP server for the MCP host to use.</p>
<h3 id="mcp-host"><a href="#mcp-host">MCP Host</a></h3>
<p>This is the application the user is actually interacting with. The AI application that coordinates and manages one or multiple MCP clients. Claude Desktop is a host. Claude Code is a host. Your custom LangGraph app would be a host.</p>
<p>The MCP client asks its server for a list of tools and resources the server provides; the server replies with a natural-language description of the capabilities of each tool and the expected format to call the tool. This information is given to the LLM; if it requires the services of one of these tools, the MCP host will instruct the relevant MCP client to call the tool. The MCP server performs the tool action and returns the results, which the MCP host then injects into the LLM conversation.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p>So the flow goes: User talks to Host → Host tells the AI what tools are available via the Client → AI decides to use a tool → Client calls the Server → Server does the thing → Result comes back to AI → AI responds to User.</p></div></div>
<p>MCP follows a client-server architecture where an MCP host establishes connections to one or more MCP servers. The MCP host accomplishes this by creating one MCP client for each MCP server. Each MCP client maintains a dedicated connection with its corresponding MCP server.</p>
<p>Local MCP servers that use the STDIO transport typically serve a single MCP client, whereas remote MCP servers that use the Streamable HTTP transport will typically serve many MCP clients.</p>
<p>Note that MCP server refers to the program that serves context data, regardless of where it runs. MCP servers can execute locally or remotely. For example, when Claude Desktop launches the filesystem server, the server runs locally on the same machine because it uses the STDIO transport. This is commonly referred to as a &quot;local&quot; MCP server. The official Sentry MCP server runs on the Sentry platform, and uses the Streamable HTTP transport. This is commonly referred to as a &quot;remote&quot; MCP server.</p>
<h2 id="what-mcp-servers-actually-expose-the-three-mcp-primitives"><a href="#what-mcp-servers-actually-expose-the-three-mcp-primitives">What MCP Servers Actually Expose (The Three MCP Primitives)</a></h2>
<p>MCP primitives are the most important concept within MCP. They define what clients and servers can offer each other. These primitives specify the types of contextual information that can be shared with AI applications and the range of actions that can be performed. An MCP server can expose three types of things.</p>
<h3 id="1-tools-actions-the-ai-can-take"><a href="#1-tools-actions-the-ai-can-take">1. Tools: Actions the AI Can Take</a></h3>
<p>Things that do something. Examples: &quot;create a Notion page&quot;, &quot;send a Slack message&quot;, &quot;query a database&quot;, &quot;book a calendar event.&quot; Tools are what the AI calls when it decides it needs to take an action. Tools are executable functions that AI applications can invoke to perform actions. They change state, interact with the world, take actions. The AI reads the tool&#x27;s name and description, decides if it needs it, and calls it.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>The mental model:</strong> A tool is like an API endpoint you have documented for the AI. You tell the AI &quot;here is what this function is called, here is what it does, here is what arguments it needs&quot;, and the AI figures out when to call it.</p></div></div>
<p>Here is a real tool definition in the JSON format MCP uses:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_1t_" class="code-block__title" title="Tool definition (JSON)">Tool definition (JSON)</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_1t_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;create_hibob_leave_request&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Submit a leave request for an employee in HiBob. Use this when an employee wants to request time off.&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;inputSchema&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;object&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;properties&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;employee_id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;The unique ID of the employee in HiBob&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;start_date&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;format&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;date&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Leave start date in YYYY-MM-DD format&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;end_date&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;format&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;date&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;leave_type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;enum&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;annual&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;sick&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;emergency&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;unpaid&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;required&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;employee_id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;start_date&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;end_date&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;leave_type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>Notice the description on the tool and on each field. These descriptions are not just documentation for humans; the AI reads them to decide if and how to use the tool. Bad descriptions mean the AI misuses or ignores your tool. Good descriptions mean the AI uses it exactly right. This is one of the most important things to get right when building MCP servers.</p>
<p><strong>The execution flow:</strong></p>
<ul>
<li><strong>Discovery:</strong> Client calls <code class="inline-code">tools/list</code> → server returns all available tool definitions.</li>
<li>AI receives those definitions in its context and understands what it can do.</li>
<li>User sends a request.</li>
<li>AI decides to use a tool → formats the call with proper arguments.</li>
<li>Client calls <code class="inline-code">tools/call</code> with tool name and arguments.</li>
<li>Server executes the actual function (calls HiBob API, writes to a database, etc.).</li>
<li>Server returns the result.</li>
<li>AI incorporates the result into its response.</li>
</ul>
<h3 id="2-resources-data-the-ai-can-read"><a href="#2-resources-data-the-ai-can-read">2. Resources: Data the AI Can Read</a></h3>
<p>Data the AI can read but cannot act on directly. Think of them as files or documents the AI can access. They are data sources that provide contextual information to AI applications. Examples: a company wiki page, a list of employees, a CSV file, API responses. Resources are read-only context. The AI does not call resources directly the way it calls tools; the app retrieves them and includes them in the conversation. Resources are passive; tools are active.</p>
<p><strong>Direct Resources</strong> have fixed URIs:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_2b_" class="code-block__title" title="Direct resources">Direct resources</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_2b_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="text" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span>calendar://events/june-2026</span></span>
<span data-line=""><span>file:///company/handbook.pdf</span></span>
<span data-line=""><span>employees://directory/engineering-team</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p><strong>Resource Templates</strong> are dynamic: URI patterns with parameters.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_2f_" class="code-block__title" title="Resource templates">Resource templates</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_2f_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="text" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span>employees://{department}/{role}</span></span>
<span data-line=""><span>→ employees://engineering/senior-engineer</span></span>
<span data-line=""><span>→ employees://finance/analyst</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>The real difference from tools in practice:</strong> If you want the AI to have company policy documents as background knowledge for a conversation, that is a resource; the app loads it and passes it to the AI before the conversation starts. If you want the AI to be able to write a new policy document, that is a tool.</p></div></div>
<h3 id="3-prompts-pre-built-templates"><a href="#3-prompts-pre-built-templates">3. Prompts: Pre-Built Templates</a></h3>
<p>Pre-built prompt templates that the server exposes. They are reusable templates that help structure interactions with language models. Less commonly used but useful for standardizing how users interact with the server. For example, a &quot;summarize this document&quot; prompt template. They are for standardizing how people interact with complex workflows.</p>
<p><strong>The practical use case:</strong> Imagine a company&#x27;s HR team uses an AI tool daily. Instead of typing a different natural language request every time and getting inconsistent results, you create a prompt called <code class="inline-code">employee_onboarding_review</code> with defined fields: employee_name, start_date, department. The HR team fills in the fields and gets a consistent, well-structured interaction every time.</p>
<p>In Claude Code or Claude Desktop, prompts typically show up as slash commands: <code class="inline-code">/plan_vacation</code>, <code class="inline-code">/onboard_employee</code>, <code class="inline-code">/generate_leave_report</code>.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p>As a concrete example, consider an MCP server that provides context about a database. It can expose tools for querying the database, a resource that contains the schema of the database, and a prompt that includes few-shot examples for interacting with the tools.</p></div></div>
<h2 id="where-mcp-fits-in-code-you-actually-write"><a href="#where-mcp-fits-in-code-you-actually-write">Where MCP Fits in Code You Actually Write</a></h2>
<p>This is where it clicks for developers. There are two situations you will encounter.</p>
<h3 id="situation-a-using-an-existing-mcp-server-someone-else-built"><a href="#situation-a-using-an-existing-mcp-server-someone-else-built">Situation A: Using an Existing MCP Server Someone Else Built</a></h3>
<p>There are already hundreds of pre-built MCP servers for popular platforms. Anthropic has shared pre-built MCP servers for popular enterprise systems like Google Drive, Slack, GitHub, Git, Postgres, and Puppeteer. Notion has one. GitHub has one. Many tools you already use have them.</p>
<p>In this situation, you do not write the server at all. You just configure your agent (or Claude Desktop/Code) to connect to it, and suddenly your agent can use those tools. This is where MCP saves the most time: you get a full Notion integration without writing a single line of Notion API code yourself.</p>
<h3 id="situation-b-building-your-own-mcp-server"><a href="#situation-b-building-your-own-mcp-server">Situation B: Building Your Own MCP Server</a></h3>
<p>When a platform does not have an MCP server yet, or you need custom functionality, you build one. For example, building MCP servers for HiBob, Entra, internal platforms, etc., so that Claude can interact with those systems. A basic MCP server in Python looks conceptually like this (simplified):</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_3b_" class="code-block__title" title="Basic MCP server (Python)">Basic MCP server (Python)</span><span class="code-block__lang">python</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_3b_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="python" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">from</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> mcp.server </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">import</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> Server</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">from</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> mcp.types </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">import</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> Tool</span></span>
<span data-line=""> </span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">server </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> Server(</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;my-hibob-server&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">)</span></span>
<span data-line=""> </span>
<span data-line=""><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB">@server.list_tools</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">()</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> def</span><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB"> list_tools</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">():</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">    return</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> [</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        Tool(</span></span>
<span data-line=""><span style="--shiki-light:#702C00;--shiki-dark:#F69D50">            name</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;get_employee&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#702C00;--shiki-dark:#F69D50">            description</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Get an employee&#x27;s details from HiBob by their ID&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#702C00;--shiki-dark:#F69D50">            inputSchema</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">                &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;object&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">                &quot;properties&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">                    &quot;employee_id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">                }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">            }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        )</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    ]</span></span>
<span data-line=""> </span>
<span data-line=""><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB">@server.call_tool</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">()</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> def</span><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB"> call_tool</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">(name: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">str</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, arguments: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">dict</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">):</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">    if</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> name </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">==</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF"> &quot;get_employee&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">:</span></span>
<span data-line=""><span style="--shiki-light:#66707B;--shiki-dark:#768390">        # call HiBob API here</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">        return</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> hibob_api.get_employee(arguments[</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;employee_id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">])</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>You define what tools exist, what they are called, what arguments they take, and what they do. Claude then discovers these tools automatically and can call them when relevant.</p>
<h3 id="mcp-in-claude-code-specifically"><a href="#mcp-in-claude-code-specifically">MCP in Claude Code Specifically</a></h3>
<p>Claude Code is where MCP becomes incredibly powerful for a developer&#x27;s day-to-day. You configure MCP servers in a config file, and then Claude Code gains access to all those tools natively: it can browse your GitHub issues, read your Notion docs, query your database, all from within your coding session.</p>
<p>For example, if you add a GitHub MCP server to Claude Code, you can say: &quot;Look at the open issues tagged &#x27;bug&#x27;, find the three most critical ones, and write fixes for them.&quot; Claude Code can actually read the issues, understand them, write the code, and if you have a Git MCP server too, even commit.</p>
<h2 id="the-two-layers-of-mcp"><a href="#the-two-layers-of-mcp">The Two Layers of MCP</a></h2>
<p>MCP consists of two layers:</p>
<ul>
<li><strong>Transport layer:</strong> Defines the communication mechanisms and channels that enable data exchange between clients and servers, including transport-specific connection establishment, message framing, and authorization.</li>
<li><strong>Data layer:</strong> Defines the JSON-RPC based protocol for client-server communication, including lifecycle management, and core primitives, such as tools, resources, prompts and notifications.</li>
</ul>
<p>Conceptually the data layer is the inner layer, while the transport layer is the outer layer.</p>
<h2 id="transport-layer-how-mcp-communicates-under-the-hood"><a href="#transport-layer-how-mcp-communicates-under-the-hood">Transport Layer: How MCP Communicates Under the Hood</a></h2>
<p>Client and server communicate using the JSON-RPC 2.0 transport protocol. There are two transport methods:</p>
<ul>
<li><strong>stdio (Standard Input/Output):</strong> For local servers running on the same machine. The client literally communicates with the server by passing messages through the terminal&#x27;s stdin/stdout. This is what you use for local development and local tools (file system access, local database, etc.).</li>
<li><strong>HTTP with SSE (Server-Sent Events):</strong> For remote servers. This is what you use when the server is hosted somewhere else, like an MCP server you deploy to a cloud service so multiple people or agents can use it. MCP uses HTTP POST for client-to-server messages with optional Server-Sent Events for streaming capabilities. This transport enables remote server communication and supports standard HTTP authentication methods including bearer tokens, API keys, and custom headers. MCP recommends using OAuth to obtain authentication tokens.</li>
</ul>
<p>The transport method (stdio vs HTTP) describes how your code communicates with the MCP server process, not how that server talks to the external API.</p>
<p>When you &quot;use the GitHub MCP server,&quot; you typically download and run that server locally on your own machine. It becomes a process running on your computer. Your MCP client talks to it via stdio (local pipes). That server then internally makes its own regular HTTP calls out to GitHub&#x27;s API. But that is invisible to MCP: from MCP&#x27;s perspective, it is just a local process.</p>
<p>So the two transport modes are not about &quot;local service vs cloud service.&quot; They are about where the MCP server process itself is running.</p>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 16rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">stdio (Standard Input/Output)</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>The MCP server runs on the same machine as your code. Your client communicates by literally writing to and reading from that process&#x27;s stdin/stdout, like piping text between programs in a terminal.
<strong>When you use it:</strong></p><ul>
<li>Running any of the pre-built MCP servers (GitHub, Notion, Google Drive, Postgres) locally on your dev machine</li>
<li>Building a server that reads local files or talks to local databases</li>
<li>Development and testing of any MCP server</li>
<li>Claude Code with locally configured MCP servers</li>
</ul></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">HTTP with SSE (Server-Sent Events)</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>The MCP server is running somewhere else: deployed to the cloud, or on a different machine on the network. Your client connects over HTTP.
<strong>When you use it:</strong></p><ul>
<li>An organization has deployed a shared MCP server that multiple engineers (or multiple Claude instances) all connect to</li>
<li>You build an MCP server and deploy it to Railway or Render so your whole team can use it</li>
<li>A company hosts a HiBob MCP server centrally that all their internal tools connect to</li>
<li>A product that exposes its capabilities as a remote MCP server (this is where the ecosystem is heading)</li>
</ul></div></li></ul>
<p>SSE specifically is how the server pushes updates to the client in real time (like notifications). HTTP POST handles client-to-server messages, and SSE handles server-to-client streaming. Together they give you full two-way communication over HTTP.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Practical rule:</strong> Start with stdio for all your projects. You will only need HTTP+SSE when you want to deploy an MCP server to the cloud for shared use.</p></div></div>
<h2 id="data-layer"><a href="#data-layer">Data Layer</a></h2>
<p>The data layer implements a JSON-RPC 2.0 based exchange protocol that defines the message structure and semantics. This layer includes:</p>
<ul>
<li><strong>Lifecycle management:</strong> Handles connection initialization, capability negotiation, and connection termination between clients and servers.</li>
<li><strong>Server features:</strong> Enables servers to provide core functionality including tools for AI actions, resources for context data, and prompts for interaction templates from and to the client.</li>
<li><strong>Client features:</strong> Enables servers to ask the client to sample from the host LLM, elicit input from the user, and log messages to the client.</li>
<li><strong>Utility features:</strong> Supports additional capabilities like notifications for real-time updates and progress tracking for long-running operations.</li>
</ul>
<h3 id="data-layer-protocol"><a href="#data-layer-protocol">Data Layer Protocol</a></h3>
<p>A core part of MCP is defining the schema and semantics between MCP clients and MCP servers. Developers will likely find the data layer, in particular the set of primitives, to be the most interesting part of MCP. It is the part of MCP that defines the ways developers can share context from MCP servers to MCP clients.</p>
<p>MCP uses JSON-RPC 2.0 as its underlying RPC protocol. Client and servers send requests to each other and respond accordingly. Notifications can be used when no response is required.</p>
<h3 id="json-rpc-20-demystified"><a href="#json-rpc-20-demystified">JSON-RPC 2.0, Demystified</a></h3>
<ul>
<li><strong>JSON</strong> is just a data format. A way of writing structured information that both humans and computers can read. Like: <code class="inline-code">{&quot;name&quot;: &quot;Jonathan&quot;, &quot;role&quot;: &quot;engineer&quot;}</code>.</li>
<li><strong>RPC</strong> stands for Remote Procedure Call. &quot;Procedure&quot; is just an old word for function. So RPC means: calling a function that lives in a different process or machine, not in your own code. Think about it: when your LangGraph agent calls a tool, it is essentially saying &quot;run this function for me.&quot; RPC is the general concept of doing that across process boundaries.</li>
<li><strong>2.0</strong> is just the version number.</li>
</ul>
<p>So JSON-RPC 2.0 is a standard that says: here is exactly how you should format a message when you want to call a function on another process, using JSON. That is it.</p>
<h3 id="what-the-messages-look-like"><a href="#what-the-messages-look-like">What the Messages Look Like</a></h3>
<p>There are three types of messages in JSON-RPC 2.0:</p>
<h3 id="a-request"><a href="#a-request">A Request</a></h3>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_57_" class="code-block__title" title="JSON-RPC request">JSON-RPC request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_57_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">3</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;tools/call&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;params&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;get_employee&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;arguments&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;employee_id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;EMP001&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<ul>
<li><code class="inline-code">jsonrpc</code>: Always &quot;2.0&quot;. Just identifies the protocol version.</li>
<li><code class="inline-code">id</code>: A unique number you assign so you can match the response to this specific request. If you send 10 requests, each gets a different id so you know which response belongs to which request.</li>
<li><code class="inline-code">method</code>: The function you want to call. In MCP, these follow a pattern like <code class="inline-code">tools/list</code>, <code class="inline-code">tools/call</code>, <code class="inline-code">resources/read</code>.</li>
<li><code class="inline-code">params</code>: The arguments you are passing to that function.</li>
</ul>
<h3 id="a-response"><a href="#a-response">A Response</a></h3>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_5d_" class="code-block__title" title="JSON-RPC response">JSON-RPC response</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_5d_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">3</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;result&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;content&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [{ </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;text&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;text&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Employee: Jonathan Mukhobe&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> }]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<ul>
<li>Same <code class="inline-code">id</code> as the request: this is how you know &quot;this response answers request #3.&quot;</li>
<li><code class="inline-code">result</code>: The returned data.</li>
</ul>
<h3 id="a-notification"><a href="#a-notification">A Notification</a></h3>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_5j_" class="code-block__title" title="JSON-RPC notification">JSON-RPC notification</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_5j_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;notifications/tools/list_changed&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>Notice that there is no <code class="inline-code">id</code> here. That is the key difference. No id means no response is expected. The sender is just saying &quot;FYI, something changed&quot; and moving on.</p>
<p>That is genuinely all JSON-RPC 2.0 is. MCP uses it as its language: every single message between an MCP client and server is one of these three formats.</p>
<h2 id="lifecycle-management-the-handshake"><a href="#lifecycle-management-the-handshake">Lifecycle Management: The Handshake</a></h2>
<p>The purpose of lifecycle management is to negotiate the capabilities that both client and server support. This section provides a step-by-step walkthrough of an MCP client-server interaction, focusing on the data layer protocol.</p>
<p>Every time your MCP client connects to a server, they go through a handshake before they do anything useful. This is lifecycle management. Think of it like two people meeting for the first time and establishing ground rules before working together. There are three phases:</p>
<h3 id="phase-1-initialize-request-client--server"><a href="#phase-1-initialize-request-client--server">Phase 1: Initialize Request (Client → Server)</a></h3>
<p>The client sends an initialize request to establish the connection and negotiate supported features. The client says: &quot;Hello, I&#x27;m connecting. Here&#x27;s what version of MCP I speak, and here&#x27;s what features I support.&quot;</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_63_" class="code-block__title" title="Initialize request">Initialize request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_63_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">1</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;initialize&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;params&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;protocolVersion&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2025-11-25&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;clientInfo&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;my-agent&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;version&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;1.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;capabilities&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;elicitation&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {}</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>The client is declaring: &quot;I support the elicitation feature&quot;, meaning if the server needs to ask the user for input, the client can handle that.</p>
<h3 id="phase-2-initialize-response-server--client"><a href="#phase-2-initialize-response-server--client">Phase 2: Initialize Response (Server → Client)</a></h3>
<p>The server responds: &quot;Understood. Here&#x27;s what version I speak, and here&#x27;s what I offer.&quot;</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_6b_" class="code-block__title" title="Initialize response">Initialize response</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_6b_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">1</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;result&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;protocolVersion&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2025-11-25&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;serverInfo&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;hibob-server&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;version&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;1.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;capabilities&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;tools&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;listChanged&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">true</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;resources&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {}</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>The server is declaring: &quot;I have tools and resources. And I can send you notifications when my tool list changes (<code class="inline-code">listChanged: true</code>).&quot;</p>
<h3 id="phase-3-initialized-notification-client--server"><a href="#phase-3-initialized-notification-client--server">Phase 3: Initialized Notification (Client → Server)</a></h3>
<p>After successful initialization, the client sends a notification to indicate it is ready. This is a notification (no id, no response expected):</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_6j_" class="code-block__title" title="Initialized notification">Initialized notification</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_6j_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;notifications/initialized&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h3 id="understanding-the-initialization-exchange"><a href="#understanding-the-initialization-exchange">Understanding the Initialization Exchange</a></h3>
<p>The initialization process is a key part of MCP&#x27;s lifecycle management and serves several critical purposes:</p>
<ul>
<li><strong>Protocol Version Negotiation:</strong> The <code class="inline-code">protocolVersion</code> field (e.g. &quot;2025-11-25&quot;) ensures both client and server are using compatible protocol versions. This prevents communication errors that could occur when different versions attempt to interact. If a mutually compatible version is not negotiated, the connection should be terminated.</li>
<li><strong>Capability Discovery:</strong> The <code class="inline-code">capabilities</code> object allows each party to declare what features they support, including which primitives they can handle (tools, resources, prompts) and whether they support features like notifications. This enables efficient communication by avoiding unsupported operations.</li>
<li><strong>Identity Exchange:</strong> The <code class="inline-code">clientInfo</code> and <code class="inline-code">serverInfo</code> objects provide identification and versioning information for debugging and compatibility purposes.</li>
</ul>
<p>In this example, the capability negotiation demonstrates how MCP primitives are declared:</p>
<p><strong>Client Capabilities:</strong></p>
<ul>
<li><code class="inline-code">&quot;elicitation&quot;: {}</code>: The client is declaring &quot;I support the elicitation feature&quot;, meaning if the server needs to ask the user for input, the client can handle that.</li>
</ul>
<p><strong>Server Capabilities:</strong></p>
<ul>
<li><code class="inline-code">&quot;tools&quot;: {&quot;listChanged&quot;: true}</code>: The server supports the tools primitive AND can send <code class="inline-code">tools/list_changed</code> notifications when its tool list changes.</li>
<li><code class="inline-code">&quot;resources&quot;: {}</code>: The server also supports the resources primitive (can handle <code class="inline-code">resources/list</code> and <code class="inline-code">resources/read</code> methods).</li>
</ul>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Why does this matter?</strong> The handshake solves a real problem: if your client and server support different versions or features, you find out immediately and can handle it gracefully rather than sending messages that neither side understands. After the handshake, both sides know exactly what they can ask of each other. From a production perspective, if a company updates a server with new features, old clients that do not support those features can still connect because they negotiated capabilities upfront. That is what makes MCP resilient at scale.</p></div></div>
<h3 id="how-this-works-in-ai-applications"><a href="#how-this-works-in-ai-applications">How This Works in AI Applications</a></h3>
<p>During initialization, the AI application&#x27;s MCP client manager establishes connections to configured servers and stores their capabilities for later use. The application uses this information to determine which servers can provide specific types of functionality (tools, resources, prompts) and whether they support real-time updates.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_7b_" class="code-block__title" title="Pseudo-code">Pseudo-code</span><span class="code-block__lang">python</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_7b_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="python" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> with</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> stdio_client(server_config) </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">as</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> (read, write):</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">    async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> with</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> ClientSession(read, write) </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">as</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session:</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        init_response </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> await</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session.initialize()</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">        if</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> init_response.capabilities.tools:</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">            app.register_mcp_server(session, </span><span style="--shiki-light:#702C00;--shiki-dark:#F69D50">supports_tools</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">True</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">)</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        app.set_server_ready(session)</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h2 id="next-step-1-tool-discovery-primitives"><a href="#next-step-1-tool-discovery-primitives">Next Step 1: Tool Discovery (Primitives)</a></h2>
<p>Now that the connection is established, the client can discover available tools by sending a <code class="inline-code">tools/list</code> request. This request is fundamental to MCP&#x27;s tool discovery mechanism: it allows clients to understand what tools are available on the server before attempting to use them.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_7h_" class="code-block__title" title="Tools list request">Tools list request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_7h_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">2</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;tools/list&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose" data-collapsed="true"><div class="code-block__header"><span id="_R_7j_" class="code-block__title" title="Tools list response">Tools list response</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_7j_" class="overflow-hidden" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">2</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;result&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;tools&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;calculator_arithmetic&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;title&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Calculator&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Perform mathematical calculations including basic arithmetic, trigonometric functions, and algebraic operations&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;inputSchema&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;object&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;properties&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">            &quot;expression&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Mathematical expression to evaluate (e.g., &#x27;2 + 3 * 4&#x27;, &#x27;sin(30)&#x27;, &#x27;sqrt(16)&#x27;)&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">            }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">          },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;required&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;expression&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      },</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;weather_current&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;title&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Weather Information&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Get current weather information for any location worldwide&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;inputSchema&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;object&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;properties&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">            &quot;location&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;City name, address, or coordinates (latitude,longitude)&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">            },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">            &quot;units&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;enum&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;metric&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;imperial&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;kelvin&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">],</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;description&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Temperature units to use in response&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">              &quot;default&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;metric&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">            }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">          },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;required&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;location&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    ]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><div class="code-block__expand"></div><span class="sr-only" aria-live="polite"></span></div></figure>
<h3 id="understanding-the-tool-discovery-exchange"><a href="#understanding-the-tool-discovery-exchange">Understanding the Tool Discovery Exchange</a></h3>
<p>The <code class="inline-code">tools/list</code> request is simple, containing no parameters. The response contains a <code class="inline-code">tools</code> array that provides comprehensive metadata about each available tool. This array-based structure allows servers to expose multiple tools simultaneously while maintaining clear boundaries between different functionalities. Each tool object in the response includes several key fields:</p>
<ul>
<li><strong>name:</strong> A unique identifier for the tool within the server&#x27;s namespace. This serves as the primary key for tool execution and should follow a clear naming pattern (e.g. <code class="inline-code">calculator_arithmetic</code> rather than just <code class="inline-code">calculate</code>).</li>
<li><strong>title:</strong> A human-readable display name for the tool that clients can show to users.</li>
<li><strong>description:</strong> Detailed explanation of what the tool does and when to use it.</li>
<li><strong>inputSchema:</strong> A JSON Schema that defines the expected input parameters, enabling type validation and providing clear documentation about required and optional parameters.</li>
</ul>
<h3 id="how-this-works-in-ai-applications-1"><a href="#how-this-works-in-ai-applications-1">How This Works in AI Applications</a></h3>
<p>The AI application fetches available tools from all connected MCP servers and combines them into a unified tool registry that the language model can access. This allows the LLM to understand what actions it can perform and automatically generates the appropriate tool calls during conversations.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_7v_" class="code-block__title" title="Pseudo-code (MCP Python SDK patterns)">Pseudo-code (MCP Python SDK patterns)</span><span class="code-block__lang">python</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_7v_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="python" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">available_tools </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> []</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">for</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">in</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> app.mcp_server_sessions():</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    tools_response </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> await</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session.list_tools()</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    available_tools.extend(tools_response.tools)</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">conversation.register_available_tools(available_tools)</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h2 id="next-step-2-tool-execution-primitives"><a href="#next-step-2-tool-execution-primitives">Next Step 2: Tool Execution (Primitives)</a></h2>
<p>The client can now execute a tool using the <code class="inline-code">tools/call</code> method. This demonstrates how MCP primitives are used in practice: after discovering available tools, the client can invoke them with appropriate arguments.</p>
<p>The <code class="inline-code">tools/call</code> request follows a structured format that ensures type safety and clear communication between client and server. Note that we are using the proper tool name from the discovery response (<code class="inline-code">weather_current</code>) rather than a simplified name:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_87_" class="code-block__title" title="Tool call request">Tool call request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_87_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">3</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;tools/call&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;params&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;weather_current&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;arguments&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;location&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;San Francisco&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;units&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;imperial&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_89_" class="code-block__title" title="Tool call response">Tool call response</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_89_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">3</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;result&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;content&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;text&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;text&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Current weather in San Francisco: 68°F, partly cloudy with light winds from the west at 8 mph. Humidity: 65%&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    ]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h3 id="key-elements-of-tool-execution"><a href="#key-elements-of-tool-execution">Key Elements of Tool Execution</a></h3>
<p>The request structure includes several important components:</p>
<ul>
<li><strong>name:</strong> Must match exactly the tool name from the discovery response (<code class="inline-code">weather_current</code>). This ensures the server can correctly identify which tool to execute.</li>
<li><strong>arguments:</strong> Contains the input parameters as defined by the tool&#x27;s inputSchema. In this example: <code class="inline-code">location</code> is &quot;San Francisco&quot; (required parameter) and <code class="inline-code">units</code> is &quot;imperial&quot; (optional parameter, defaults to &quot;metric&quot; if not specified).</li>
<li><strong>JSON-RPC Structure:</strong> Uses standard JSON-RPC 2.0 format with unique id for request-response correlation.</li>
</ul>
<h3 id="understanding-the-tool-execution-response"><a href="#understanding-the-tool-execution-response">Understanding the Tool Execution Response</a></h3>
<p>The response demonstrates MCP&#x27;s flexible content system:</p>
<ul>
<li><strong>content Array:</strong> Tool responses return an array of content objects, allowing for rich, multi-format responses (text, images, resources, etc.).</li>
<li><strong>Content Types:</strong> Each content object has a type field. In this example, <code class="inline-code">&quot;type&quot;: &quot;text&quot;</code> indicates plain text content, but MCP supports various content types for different use cases.</li>
<li><strong>Structured Output:</strong> The response provides actionable information that the AI application can use as context for language model interactions.</li>
</ul>
<p>This execution pattern allows AI applications to dynamically invoke server functionality and receive structured responses that can be integrated into conversations with language models.</p>
<h3 id="how-this-works-in-ai-applications-2"><a href="#how-this-works-in-ai-applications-2">How This Works in AI Applications</a></h3>
<p>When the language model decides to use a tool during a conversation, the AI application intercepts the tool call, routes it to the appropriate MCP server, executes it, and returns the results back to the LLM as part of the conversation flow. This enables the LLM to access real-time data and perform actions in the external world.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_8t_" class="code-block__title" title="Pseudo-code for tool execution">Pseudo-code for tool execution</span><span class="code-block__lang">python</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_8t_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="python" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> def</span><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB"> handle_tool_call</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">(conversation, tool_name, arguments):</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    session </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> app.find_mcp_session_for_tool(tool_name)</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    result </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> await</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session.call_tool(tool_name, arguments)</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    conversation.add_tool_result(result.content)</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h2 id="next-step-3-real-time-updates-notifications"><a href="#next-step-3-real-time-updates-notifications">Next Step 3: Real-Time Updates (Notifications)</a></h2>
<p>MCP supports real-time notifications that enable servers to inform clients about changes without being explicitly requested. This demonstrates the notification system, a key feature that keeps MCP connections synchronized and responsive.</p>
<p>When the server&#x27;s available tools change, such as when new functionality becomes available, existing tools are modified, or tools become temporarily unavailable, the server can proactively notify connected clients:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_95_" class="code-block__title" title="Tool list change notification">Tool list change notification</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_95_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;notifications/tools/list_changed&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h3 id="key-features-of-mcp-notifications"><a href="#key-features-of-mcp-notifications">Key Features of MCP Notifications</a></h3>
<ul>
<li><strong>No Response Required:</strong> Notice there is no id field in the notification. This follows JSON-RPC 2.0 notification semantics where no response is expected or sent.</li>
<li><strong>Capability-Based:</strong> This notification is only sent by servers that declared <code class="inline-code">&quot;listChanged&quot;: true</code> in their tools capability during initialization.</li>
<li><strong>Event-Driven:</strong> The server decides when to send notifications based on internal state changes, making MCP connections dynamic and responsive.</li>
</ul>
<h3 id="client-response-to-notifications"><a href="#client-response-to-notifications">Client Response to Notifications</a></h3>
<p>Upon receiving this notification, the client typically reacts by requesting the updated tool list. This creates a refresh cycle that keeps the client&#x27;s understanding of available tools current:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_9f_" class="code-block__title" title="Refreshed tools list request">Refreshed tools list request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_9f_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;jsonrpc&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;2.0&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;id&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#023B95;--shiki-dark:#6CB6FF">4</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;tools/list&quot;</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h3 id="why-notifications-matter"><a href="#why-notifications-matter">Why Notifications Matter</a></h3>
<p>This notification system is crucial for several reasons:</p>
<ul>
<li><strong>Dynamic Environments:</strong> Tools may come and go based on server state, external dependencies, or user permissions.</li>
<li><strong>Efficiency:</strong> Clients do not need to poll for changes; they are notified when updates occur.</li>
<li><strong>Consistency:</strong> Ensures clients always have accurate information about available server capabilities.</li>
<li><strong>Real-time Collaboration:</strong> Enables responsive AI applications that can adapt to changing contexts.</li>
</ul>
<p>This notification pattern extends beyond tools to other MCP primitives, enabling comprehensive real-time synchronization between clients and servers.</p>
<h3 id="how-this-works-in-ai-applications-3"><a href="#how-this-works-in-ai-applications-3">How This Works in AI Applications</a></h3>
<p>When the AI application receives a notification about changed tools, it immediately refreshes its tool registry and updates the LLM&#x27;s available capabilities. This ensures that ongoing conversations always have access to the most current set of tools, and the LLM can dynamically adapt to new functionality as it becomes available.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_9t_" class="code-block__title" title="Pseudo-code for notification handling">Pseudo-code for notification handling</span><span class="code-block__lang">python</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_9t_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="python" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">async</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> def</span><span style="--shiki-light:#622CBC;--shiki-dark:#DCBDFB"> handle_tools_changed_notification</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">(session):</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    tools_response </span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">=</span><span style="--shiki-light:#A0111F;--shiki-dark:#F47067"> await</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> session.list_tools()</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    app.update_available_tools(session, tools_response.tools)</span></span>
<span data-line=""><span style="--shiki-light:#A0111F;--shiki-dark:#F47067">    if</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> app.conversation.is_active():</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        app.conversation.notify_llm_of_new_capabilities()</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<h2 id="client-features-what-the-client-offers-the-server"><a href="#client-features-what-the-client-offers-the-server">Client Features: What the Client Offers the Server</a></h2>
<p>This is the part most people miss. MCP is two-way. Servers expose tools and resources to clients, but clients also expose features to servers. The server can use these features to build richer, more powerful interactions.</p>
<h3 id="elicitation-the-server-asks-the-user-a-question"><a href="#elicitation-the-server-asks-the-user-a-question">Elicitation: The Server Asks the User a Question</a></h3>
<p>Normally: User → Client → AI → Server (the server just receives and acts). With elicitation, the server can pause its work and say to the client: &quot;I need more information from the user before I can continue.&quot;</p>
<p>A concrete example: Your agent is in the middle of booking a flight. The MCP server handling the booking has found options and is ready to confirm. But before finalizing, it needs the user&#x27;s passport number and seat preference. The server can use elicitation to pause and ask:</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_a9_" class="code-block__title" title="Elicitation request">Elicitation request</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_a9_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;method&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;elicitation/create&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;params&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;message&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Before I confirm your booking, I need a few details:&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">    &quot;requestedSchema&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;object&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;properties&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;passport_number&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;seat_preference&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: {</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;string&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">,</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">          &quot;enum&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;window&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;aisle&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;middle&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">        },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">        &quot;confirm_booking&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;type&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;boolean&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">      },</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">      &quot;required&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;passport_number&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;confirm_booking&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p>The client shows this as a UI form to the user, collects the structured response, and sends it back to the server. The server then continues with that information.</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Why this matters for production agents:</strong> Without elicitation, if the server needs more info, it has to either fail or make assumptions. With elicitation, you can build agents that handle missing information gracefully mid-workflow, a critical feature for real enterprise automations where you cannot always predict what info you will need upfront.</p></div></div>
<h3 id="roots-filesystem-scope-boundaries"><a href="#roots-filesystem-scope-boundaries">Roots: Filesystem Scope Boundaries</a></h3>
<p>Roots are how the client tells the server which directories it is allowed to operate in. If you connect a filesystem MCP server to help with a project, you do not want it looking through your entire machine; you point it at specific folders.</p>
<figure data-rehype-pretty-code-figure=""><div class="code-block not-prose"><div class="code-block__header"><span id="_R_aj_" class="code-block__title" title="Roots declaration">Roots declaration</span><span class="code-block__lang">json</span><div class="code-block__actions"></div></div><pre tabindex="0" aria-labelledby="_R_aj_" class="" data-theme="github-light-high-contrast github-dark-dimmed"><code data-language="json" data-theme="github-light-high-contrast github-dark-dimmed" style="display:grid"><span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">{</span></span>
<span data-line=""><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">  &quot;roots&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: [</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;uri&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;file:///home/jonathan/projects/projects&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Projects&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> },</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">    { </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;uri&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;file:///home/jonathan/docs/research&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">, </span><span style="--shiki-light:#024C1A;--shiki-dark:#8DDB8C">&quot;name&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">: </span><span style="--shiki-light:#032563;--shiki-dark:#96D0FF">&quot;Research&quot;</span><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7"> }</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">  ]</span></span>
<span data-line=""><span style="--shiki-light:#0E1116;--shiki-dark:#ADBAC7">}</span></span></code></pre><span class="sr-only" aria-live="polite"></span></div></figure>
<p><strong>Important caveat from the docs:</strong> Roots are a convention, not a hard security boundary. A well-behaved server respects them. But they do not enforce access at the OS level. Think of them as telling a trusted contractor &quot;please only work in these rooms&quot;; they should follow that, but it is not a locked door. Real security enforcement happens through OS-level file permissions, not MCP roots.</p>
<p>Claude Code uses roots automatically when you open a folder: it tells any connected MCP servers &quot;this is the project scope.&quot;</p>
<h3 id="sampling-the-server-asks-the-ai-to-think"><a href="#sampling-the-server-asks-the-ai-to-think">Sampling: The Server Asks the AI to Think</a></h3>
<p>This is the most advanced and counterintuitive client feature. Sampling allows the server to make a request back to the AI (through the client) during its own operation.</p>
<p>Why would a server need to ask the AI? Sometimes a tool needs intelligence to decide what to do. Example: your MCP server retrieves 47 flight options. It could return all 47 to the AI in the main conversation, but that would flood the context window. Instead, the server can use sampling: &quot;Hey client, can you ask the AI to look at these 47 flights and tell me the top 3 for this user&#x27;s preferences?&quot; The AI analyzes them in a separate call, returns a structured answer, and the server uses that to give a clean response.</p>
<p><strong>Why this keeps the client in control:</strong> The sampling request goes through the client, and the client can show it to the user for approval before forwarding it to the AI. The server cannot make secret AI calls on the user&#x27;s behalf. The client is the gatekeeper.</p>
<p><strong>For your projects:</strong> Sampling is advanced and you likely will not need it in your first MCP builds. But knowing it exists helps you understand why MCP is powerful: you can build server-side logic that uses AI reasoning without the server needing its own LLM integration.</p>
<h2 id="notifications-one-way-messages"><a href="#notifications-one-way-messages">Notifications: One-Way Messages</a></h2>
<p>You have seen these throughout. The pattern is simple: a notification is a JSON-RPC message with no <code class="inline-code">id</code> field. This means no response is expected. It is a signal, not a request. The most important notifications in practice:</p>
<ul>
<li><code class="inline-code">notifications/tools/list_changed</code>: Server tells the client its tool list has changed. The client should call <code class="inline-code">tools/list</code> again to get the updated list. This matters when servers dynamically add or remove tools based on user permissions or system state.</li>
<li><code class="inline-code">notifications/initialized</code>: Client tells the server it is ready to work (sent after the initialize handshake).</li>
<li><code class="inline-code">notifications/progress</code>: For long-running operations. The server sends progress updates: &quot;Step 1 of 5 complete... Step 2 of 5...&quot;</li>
</ul>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Why notifications matter at large company scale:</strong> Imagine 2,500 employees all using an agent connected to an MCP server. If that server updates (new tools added, permissions changed for certain users), every connected client needs to know. Without notifications, clients would be working with stale information. With notifications, all 2,500 active connections get updated in real time without each one having to poll the server constantly.</p></div></div>
<h2 id="tasks-experimental-long-running-operations"><a href="#tasks-experimental-long-running-operations">Tasks (Experimental): Long-Running Operations</a></h2>
<p>This is a newer, still-experimental feature. Normal tool calls are synchronous: you call a tool, you wait, you get the result. But some operations take minutes or hours, like &quot;run a batch report on all 7 million customers&quot; or &quot;process this large dataset.&quot; Tasks are durable execution wrappers that let you fire off a long operation and check on it later:</p>
<ul>
<li>Call a tool that starts a task → get back a <code class="inline-code">task_id</code> immediately</li>
<li>Do other things</li>
<li>Later, check <code class="inline-code">task_id</code> for status: &quot;running&quot; / &quot;complete&quot; / &quot;failed&quot;</li>
<li>When complete, retrieve the result</li>
</ul>
<p>This is analogous to Celery in Django: you enqueue a job, it runs in the background, you check on it. Same concept, just standardized in MCP. For large company automations that touch large data sets or multi-step workflows, this is the mechanism that would make those agents non-blocking and production-grade.</p>
<h2 id="versioning-how-mcp-manages-change"><a href="#versioning-how-mcp-manages-change">Versioning: How MCP Manages Change</a></h2>
<h3 id="the-date-based-version-format-yyyy-mm-dd"><a href="#the-date-based-version-format-yyyy-mm-dd">The Date-Based Version Format (YYYY-MM-DD)</a></h3>
<p>MCP uses dates as version numbers, not v1.0 or v2.0. The current version as of this article is 2025-11-25. The date represents the last time a breaking (backwards-incompatible) change was made.</p>
<p>The key rule: If Anthropic adds a new feature to MCP but does not break anything that already works, the version date stays the same. They only bump the date when they make a change that would break existing implementations. This is a deliberately conservative approach: the protocol can keep improving without forcing everyone to upgrade constantly.</p>
<h3 id="individual-feature-states"><a href="#individual-feature-states">Individual Feature States</a></h3>
<p>Beyond whole-version states, individual features within a version can have their own lifecycle:</p>
<ul>
<li><strong>Deprecated:</strong> The feature still works but is being phased out. They give at least 12 months warning (90 days minimum in exceptional cases) before removing it. Use this time to migrate away.</li>
<li><strong>Removed:</strong> Gone in a new version. You will see it in older specs but it will not work in current servers.</li>
</ul>
<h3 id="version-negotiation-during-initialization"><a href="#version-negotiation-during-initialization">Version Negotiation During Initialization</a></h3>
<p>This is the practical part. During the initialization handshake (Phase 1 above), the client says &quot;I speak version 2025-11-25.&quot; The server responds with what version it uses. If they are on different versions:</p>
<ul>
<li>If backwards compatible → they can still work together</li>
<li>If incompatible → the connection terminates gracefully with an error, rather than silently failing or producing weird behavior</li>
</ul>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Why this matters when you build:</strong> Always check the current protocol version before building. Build against &quot;Current,&quot; not &quot;Draft.&quot; If you ever see &quot;Deprecated&quot; next to a feature you are using, plan your migration: you have at least 12 months, but do not leave it to the last minute.</p></div></div>
<p>I rebuilt my own WhatsApp agent on this protocol, and the <a href="https://www.jonathanmuk.com/projects/kylie-mcp">Kylie MCP case study</a> shows what each of these pieces looks like in a real repository.</p>
<h2 id="what-you-actually-write-in-summary"><a href="#what-you-actually-write-in-summary">What You Actually Write, In Summary</a></h2>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 13rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">When You Build an MCP Server</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>You write tool definitions (name, description, inputSchema), tool handlers (the actual function that runs), resource endpoints (if exposing data), and optionally notification senders (if your tool list changes dynamically).</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">When You Build an MCP Client (Your Agent)</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>You write the initialization handshake, the <code class="inline-code">tools/list</code> call on startup, tool call routing when the AI decides to use a tool, notification listeners to keep your tool registry fresh, and optionally elicitation UI if you want users to answer mid-task questions.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">When You Deploy to Production</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>You move from stdio (local) to HTTP+SSE (remote), add proper authentication (OAuth or API keys in the HTTP headers), add observability (logging on each tool call, the same discipline as an Opik-style tracing setup), and version pin to the Current spec.</p></div></li></ul>
<p>You are now ready to go build. The jump from understanding to hands-on will make everything click, especially the initialization sequence, which looks abstract until you see it happen in logs for the first time.</p>
<blockquote>
<p>Think of MCP like a USB-C port for AI applications: one standard connector between any AI application and any external system, so nobody has to keep soldering custom wires.</p>
</blockquote>]]></content:encoded>
            <author>mukjonas256@gmail.com (Jonathan Mukhobe)</author>
            <category domain="https://www.jonathanmuk.com/insights?tag=MCP">MCP</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=AI%20Agents">AI Agents</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=JSON-RPC">JSON-RPC</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Claude">Claude</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=AI%20Engineering">AI Engineering</category>
            <enclosure length="114462" type="image/jpeg" url="https://www.jonathanmuk.com/og/insights/model-context-protocol.jpg"/>
        </item>
        <item>
            <title><![CDATA[RAG Simplified: How to Give Your LLM a Long-Term Memory]]></title>
            <link>https://www.jonathanmuk.com/insights/rag-simplified</link>
            <guid isPermaLink="false">https://www.jonathanmuk.com/insights/rag-simplified</guid>
            <pubDate>Thu, 17 Jul 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[Why fine-tuning is often the wrong answer, how vector databases work, and the end-to-end retrieval augmented generation flow from tokens to grounded reply.]]></description>
            <content:encoded><![CDATA[<h2 id="retrieval-augmented-generation-rag"><a href="#retrieval-augmented-generation-rag">Retrieval Augmented Generation (RAG)</a></h2>
<p>RAG is a technique that exerts influence on a model&#x27;s response through the integration of our own data. Rather than baking knowledge directly into model weights, it retrieves the most relevant information at query time and injects it into the prompt, giving the model an accurate, up-to-date context without ever retraining.</p>
<h2 id="why-fine-tuning-is-often-the-wrong-answer"><a href="#why-fine-tuning-is-often-the-wrong-answer">Why Fine-Tuning Is Often the Wrong Answer</a></h2>
<p>One possible way to influence a model with your own data is through fine-tuning and incorporating it into the training process. This approach enables the model to construct responses taking into account the provided data. However, often this may not be the optimal solution for various reasons:</p>
<ul>
<li>Cost and effort involved in fine-tuning a model, including computer resources and GPU access.</li>
<li>Information is not valid forever. It might need to be changed, which calls for fine-tuning every time the information is modified.</li>
<li>There is no mechanism to make the model forget what it has learned.</li>
<li>In a scenario where information increases periodically, fine-tuning is not the best approach.</li>
<li>A model can receive data in two ways: through fine-tuning and through the system prompt. But the information we can give to the prompt is limited.</li>
</ul>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p><strong>Solution:</strong> We can use RAG such that the model receives necessary information for the specific question the user provides, pulled from a knowledge base at query time.</p></div></div>
<h2 id="steps-followed-by-a-rag-system"><a href="#steps-followed-by-a-rag-system">Steps Followed by a RAG System</a></h2>
<h3 id="knowledge-storage"><a href="#knowledge-storage">Knowledge Storage</a></h3>
<p>The first step is to store the information the model can draw upon. This knowledge base can include structured data (like databases), unstructured data (like text documents), or even code. We use a vector database to store this information.</p>
<h3 id="question-reception"><a href="#question-reception">Question Reception</a></h3>
<p>The RAG system receives the user&#x27;s question or request. This could be in natural language, just like a question you would ask a person.</p>
<h3 id="information-retrieval"><a href="#information-retrieval">Information Retrieval</a></h3>
<p>This is where the &quot;Retrieval&quot; in RAG comes into play. The system selects relevant information to the user&#x27;s question from the stored data. In a vector database, this involves vector comparison to retrieve related information.</p>
<p><em>Vector Comparison</em> is how the system figures out which stored information is most similar to your question, using math behind the scenes to compare meaning. It matches the meanings of your question to the most relevant stored documents, not by keywords, but by how close their meanings are in a mathematical space.</p>
<h3 id="prompt-construction"><a href="#prompt-construction">Prompt Construction</a></h3>
<p>A prompt is constructed using the relevant information and the user&#x27;s question. This prompt is designed to effectively communicate the context of the user&#x27;s query to the language model. In this way, not only is the model provided with information to construct its response, but hallucinations are also reduced since the model does not have to invent information it does not know.</p>
<h3 id="calling-the-model"><a href="#calling-the-model">Calling the Model</a></h3>
<p>The prompt is passed to the language model. The model uses the context provided in the prompt to generate a response.</p>
<h3 id="response-from-the-model"><a href="#response-from-the-model">Response from the Model</a></h3>
<p>The model&#x27;s response is presented to the user. If the RAG system has been constructed correctly, this response should be relevant and grounded in the retrieved information from the knowledge base.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img1.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img1.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img1.png" alt="RAG system flow diagram: Question to Vector DB to Powered Prompt to Model" width="891" height="426" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">RAG system flow diagram: Question to Vector DB to Powered Prompt to Model</figcaption></figure>
<h2 id="how-vector-databases-work"><a href="#how-vector-databases-work">How Vector Databases Work</a></h2>
<p>Of all the steps mentioned, one is particularly important and sensitive: searching for relevant information to answer the user&#x27;s question. This is where vector databases come into play, assisting in selecting which information from the database is relevant to what the user asks.</p>
<p>Vector databases store information as numerical representations called vectors. So it is necessary to transform text into vectors before storing it.</p>
<h3 id="tokenization"><a href="#tokenization">Tokenization</a></h3>
<p>The first step is to tokenize the text. Tokens are the smallest units of text that are meaningful to the model. There are various types of tokens, including words, subwords, characters, and byte-pair encodings. The choice of token type depends on the specific use case and the language being modeled.</p>
<h3 id="vectorization"><a href="#vectorization">Vectorization</a></h3>
<p>Tokens are converted into vectors, which are numerical representations of objects in a continuous vector space. They capture the semantic meaning or properties of those objects in a way that can be efficiently searched. A vector represents a point in a multidimensional space, anywhere from roughly 100 to 4,000 dimensions in practice.</p>
<h3 id="embeddings-and-semantic-proximity"><a href="#embeddings-and-semantic-proximity">Embeddings and Semantic Proximity</a></h3>
<p>The vector, also known as an embedding, captures the semantic meaning of the stored text. The magic lies in how these embeddings are generated: words with similar meanings have vectors that are closer together in this multidimensional space than words with different meanings.</p>
<p>For example, &quot;cat&quot; and &quot;kitty&quot; have very similar meanings, even though the words themselves are quite different letter by letter. Their embeddings will be numerically close because the model has learned they appear in similar contexts.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img2.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img2.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img2.png" alt="Tokenization and neural network vectorization diagram (cat example)" width="817" height="249" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Tokenization and neural network vectorization diagram (cat example)</figcaption></figure>
<h3 id="distance-calculation-and-search"><a href="#distance-calculation-and-search">Distance Calculation and Search</a></h3>
<p>Once text is converted, it is possible to calculate the difference between one vector and another, or search for vectors that are closer to a specific one. That is how it is possible to find similar texts to a reference one. Mathematically, there is not much difference between calculating the distance between two points whether they are in two, three, or any number of dimensions.</p>
<h2 id="where-do-embedding-numbers-come-from"><a href="#where-do-embedding-numbers-come-from">Where Do Embedding Numbers Come From?</a></h2>
<p>The numbers in vector embeddings do not just appear by chance. They come from machine learning models that have been trained on huge amounts of text (or other types of data like images, audio, etc.).</p>
<div role="note" class="callout my-8 rounded-md border border-line bg-page-subtle px-5 py-4" aria-label="Note" data-variant="note"><p class="mb-2 flex items-center gap-2 font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Note</p><div class="text-[0.95em] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-3"><p>Think of it like teaching a child through examples. Suppose you give a child thousands of examples of how words are used in real sentences: &quot;The cat climbed the tree,&quot; &quot;The kitten purred softly,&quot; &quot;A lion is a big cat.&quot; Eventually the child starts to understand relationships between words. This is basically what deep learning models do: they read millions of examples, learn patterns, and turn each word, sentence, or document into a vector of numbers that captures those learned relationships.</p></div></div>
<h3 id="behind-the-scenes-neural-networks"><a href="#behind-the-scenes-neural-networks">Behind the Scenes: Neural Networks</a></h3>
<p>These embeddings are created by models like:</p>
<ul>
<li><strong>Word2Vec</strong></li>
<li><strong>GloVe</strong></li>
<li><strong>BERT</strong></li>
<li><strong>OpenAI&#x27;s models</strong>, etc.</li>
</ul>
<p>These models use neural networks to analyze text contextually and assign each word (or sentence) a place in a high-dimensional space. So instead of saying &quot;cat and dog are related because they appear together,&quot; they say &quot;let me place cat and dog in a vector space so they end up close together because the model has seen them used in similar ways.&quot;</p>
<p>When you hear about embedding models or vectorization tools like OpenAI Embeddings API, Hugging Face Transformers, SentenceTransformers, or spaCy, these are tools that generate those vectors from text, making it ready to store and search in a vector database.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img3.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img3.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img3.png" alt="3D vector space visualization showing Wolf, Dog, Cat, Apple, Banana cluster positions" width="966" height="622" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">3D vector space visualization showing Wolf, Dog, Cat, Apple, Banana cluster positions</figcaption></figure>
<h2 id="summary-end-to-end-rag-flow"><a href="#summary-end-to-end-rag-flow">Summary: End-to-End RAG Flow</a></h2>
<p>In summary, this is how Retrieval Augmented Generation works:</p>
<ul>
<li>The text to store is tokenized and converted into embeddings, which are then stored in the vector database.</li>
<li>The user&#x27;s question is also tokenized and converted into an embedding.</li>
<li>The vector database is queried to find the embeddings that are closest to the user&#x27;s question embedding.</li>
<li>These embeddings are converted back into text and returned as context for the model.</li>
</ul>
<p>Here is a brief summary of the process text undergoes before reaching a language model. First it is tokenized, meaning it is divided into small parts. Then it is converted into vectors, which in the world of large language models are known as embeddings. This vector, containing numbers, is passed to the model, and it generates another vector as output. The output vector undergoes the reverse process and is converted into text, which is what we see as a response.</p>
<p>The key is to utilize selected text related to the question and build a powered prompt with the information and the user request. In this way, the model has the necessary information to build a proper response. To select the text to use to enrich the prompt, a vectorized database is employed, enabling a search based on vector similarity. This ensures that the model has information relevant to what it needs to respond.</p>
<blockquote>
<p>The important thing is to understand that embeddings are numerical representations capable of capturing the meaning of sentences, and that they allow you to perform simple vector operations such as calculating the distance between them.</p>
</blockquote>]]></content:encoded>
            <author>mukjonas256@gmail.com (Jonathan Mukhobe)</author>
            <category domain="https://www.jonathanmuk.com/insights?tag=RAG">RAG</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Vector%20Databases">Vector Databases</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=LLM">LLM</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Embeddings">Embeddings</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=AI%20Engineering">AI Engineering</category>
            <enclosure length="128619" type="image/jpeg" url="https://www.jonathanmuk.com/og/insights/rag-simplified.jpg"/>
        </item>
        <item>
            <title><![CDATA[Load Balancing at Scale: Strategies That Keep Systems Alive]]></title>
            <link>https://www.jonathanmuk.com/insights/load-balancing</link>
            <guid isPermaLink="false">https://www.jonathanmuk.com/insights/load-balancing</guid>
            <pubDate>Wed, 04 Jun 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[A practical guide to load balancing: hardware, software and cloud balancers, OSI layer classification, six routing algorithms and the metrics worth monitoring.]]></description>
            <content:encoded><![CDATA[<h2 id="what-is-load-balancing"><a href="#what-is-load-balancing">What Is Load Balancing?</a></h2>
<p>At its core, load balancing is the practice of distributing incoming network traffic across multiple servers so that no single server becomes overwhelmed. Think of it like a supermarket with multiple checkout lanes: instead of everyone queuing at one cashier, a manager directs customers to whichever lane is shortest. The customers get served faster, and no cashier burns out.</p>
<p>For example, when a user hits a React app and then a Django API, something has to decide which server instance handles that request. That &quot;something&quot; is a load balancer.</p>
<p><strong>Why it matters:</strong> A single server has a ceiling, in terms of CPU, RAM, and network bandwidth. Load balancers let you scale horizontally (add more servers) rather than just vertically (buy a bigger server). They also give you redundancy: if one server crashes, traffic is automatically rerouted to the healthy ones.</p>
<h2 id="types-of-load-balancers"><a href="#types-of-load-balancers">Types of Load Balancers</a></h2>
<p>There are three broad categories, each with different trade-offs in cost, control, and capability.</p>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 13rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Hardware Load Balancers</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Physical appliances from vendors like F5 or Citrix. They are extremely fast (purpose-built silicon), handle millions of connections, and are found in large banks and telecoms. The downside: they are expensive (tens of thousands of dollars), inflexible, and hard to scale. As a developer building web apps, you will almost certainly never touch one of these.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Software Load Balancers</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Programs running on commodity hardware or VMs. NGINX, HAProxy, and Envoy are the big names. You get full control, they are cheap, and they are extraordinarily configurable. The trade-off is that you are responsible for running and scaling the software itself.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Cloud-Based Load Balancers</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Managed services like AWS ALB, Cloudflare Load Balancing, and GCP Cloud Load Balancing. You configure them via a dashboard or API and the provider handles everything else. This is the most relevant category for a modern web developer stack.</p></div></li></ul>
<h2 id="load-balancer-classification-by-osi-layer"><a href="#load-balancer-classification-by-osi-layer">Load Balancer Classification by OSI Layer</a></h2>
<p>Load balancers do not all operate at the same level of the network stack, and the layer they operate at determines what information they can use to make routing decisions.</p>
<h3 id="layer-4-transport"><a href="#layer-4-transport">Layer 4 (Transport)</a></h3>
<p>The load balancer sees the IP address, port, and whether it is a TCP or UDP connection, but it never opens the packet and reads what is inside. Think of it like a postal sorting facility that routes parcels by zip code, without opening them. It is extremely fast because the inspection is minimal.</p>
<p><strong>Good for:</strong> raw TCP traffic, non-HTTP protocols (like databases, game servers, MQTT), and situations where you need maximum throughput.</p>
<h3 id="layer-7-application"><a href="#layer-7-application">Layer 7 (Application)</a></h3>
<p>The load balancer actually reads the HTTP request: the URL, headers, cookies, even the body. Now it can make smart decisions. Route <code class="inline-code">/api/*</code> requests to your Django servers, <code class="inline-code">/</code> to your React build, and <code class="inline-code">/uploads/*</code> to object storage. Route requests with a specific session cookie to the same server that created that session. This is what Cloudflare, Vercel, and most modern cloud load balancers do.</p>
<h3 id="global-server-load-balancing-gslb"><a href="#global-server-load-balancing-gslb">Global Server Load Balancing (GSLB)</a></h3>
<p>This operates above both L4 and L7. Instead of routing requests between servers in one location, a GSLB routes users to the right data center (or region) before any L4/L7 load balancing even begins. Two mechanisms power this:</p>
<ul>
<li><strong>DNS-based routing:</strong> When a user queries a domain, the DNS server returns a different IP based on where the user is. A user in Kampala gets the IP of the closest data center; a user in Berlin gets a different one. Cloudflare&#x27;s Load Balancing product does exactly this.</li>
<li><strong>Anycast networking:</strong> Multiple servers worldwide advertise the same IP address. Internet routing protocols (BGP) automatically direct each user&#x27;s traffic to the topologically nearest server. Cloudflare&#x27;s entire network runs on Anycast, which is why Cloudflare is so fast globally.</li>
</ul>
<h2 id="load-balancing-algorithms"><a href="#load-balancing-algorithms">Load Balancing Algorithms</a></h2>
<p>Once traffic arrives at a load balancer, it needs to decide which server gets each request. Here are the main algorithms:</p>
<h3 id="round-robin"><a href="#round-robin">Round Robin</a></h3>
<p>Here, servers get requests in a sequential order i.e the first request goes to server 1, then the next request to server 2, and the third to server 3, then the next request back to server 1, so requests go to servers in a fixed cycle: 1, 2, 3, 1, 2, 3. It works when all your servers are identical(have the same capabilities) and requests take similar time. Simple and predictable.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img6.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img6.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img6.png" alt="Round Robin diagram" width="944" height="484" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Round Robin diagram</figcaption></figure>
<h3 id="sticky-round-robin-sessions"><a href="#sticky-round-robin-sessions">Sticky Round Robin (Sessions)</a></h3>
<p>Same as round robin, but once a client is assigned a server, all their requests go to that same server via a cookie or IP. Great for stateful apps that store sessions in memory. If your Django app stores session data in memory (not Redis), a user who gets routed to Server 2 for login must keep going to Server 2, or their session disappears. Sticky sessions handle this, though the better long-term solution is externalizing sessions to Redis.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img7.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img7.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img7.png" alt="Sticky Round Robin diagram" width="909" height="483" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Sticky Round Robin diagram</figcaption></figure>
<h3 id="ipurl-hash-deterministic"><a href="#ipurl-hash-deterministic">IP/URL Hash (Deterministic)</a></h3>
<p>A hash of the client&#x27;s IP (or the URL) determines the server. This algorithm determines which server gets a request based on the hash of the client&#x27;s IP address, for example, if client 1 makes a request to the load balancer, the load balancer uses the client&#x27;s IP address, hashes it and sends the request to an appropriate server, and all future requests made by that client are redirected to tha same server by the load balancer using the IP hashing algorithm. This is useful when you want the client to consistently connect to the same server.The same input always maps to the same server. Good for caching (cache locality stays on one server), but problematic if one IP generates massive traffic (such as a mobile carrier NAT-ing thousands of users behind one IP).</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img9.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img9.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img9.png" alt="IP Hash routing diagram" width="870" height="473" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">IP Hash routing diagram</figcaption></figure>
<h3 id="least-connections-dynamic"><a href="#least-connections-dynamic">Least Connections (Dynamic)</a></h3>
<p>Each new request goes to whichever server currently has the fewest active connections. Like joining the shortest queue at a theme park ride. Excellent for requests with variable processing time. If some requests take 5ms and others take 5 seconds, you do not want servers getting buried while others sit idle. This algorithm prevents that.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img10.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img10.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img10.png" alt="Least Connections dynamic routing diagram" width="949" height="495" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Least Connections dynamic routing diagram</figcaption></figure>
<h3 id="least-response-time-most-responsive"><a href="#least-response-time-most-responsive">Least Response Time (Most Responsive)</a></h3>
<p>This algorithm is more focused on responsiveness. The load balancer chooses the server with the lowest response time, and fewest active connections. So, it routes to the server with the lowest combination of active connections and fastest average response time. So, first it&#x27;ll try sending as many requests as possible to the most responsive server, but it also takes into account the number of active connections. If Server 3 has slightly more connections but responds in 8ms versus Server 1&#x27;s 45ms, Server 3 still wins. This is effective when the goal is to provide the fastest response time to requests, and you have different servers with different capabilities.The most intelligent option, used by NGINX Plus and Cloudflare.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img11.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img11.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img11.png" alt="Least Response Time algorithm diagram" width="833" height="476" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Least Response Time algorithm diagram</figcaption></figure>
<h3 id="weighted-algorithm"><a href="#weighted-algorithm">Weighted Algorithm</a></h3>
<p>These are variants of any of the algorithms; round robin, least connections, least response time, IP hash, that can also be weighted for example weighted round robin, weighted least connections. In this case, servers are assigned weights based on their capacity and performance metrics. The load balancer takes that into account when redirecting traffic. First, it&#x27;ll try sending as many requests to the server with the highest weight and then followed by the other servers if the one with the highest weight gets overwhelmed. A server with weight 3 gets 3 times more requests than one with weight 1. Like assigning more tables in a restaurant to experienced waiters. Use this when your servers have unequal specs. If you have a 4-core server and an 8-core server, give the 8-core double the weight.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><picture><source type="image/avif" srcSet="/figures/img8.720.avif 720w" sizes="(min-width: 64rem) 720px, 100vw"/><source type="image/webp" srcSet="/figures/img8.720.webp 720w" sizes="(min-width: 64rem) 720px, 100vw"/><img src="https://www.jonathanmuk.com/figures/img8.png" alt="Weighted Round Robin diagram" width="849" height="497" loading="lazy" decoding="async" class="h-auto w-full"/></picture></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Weighted Round Robin diagram</figcaption></figure>
<p>Load balancing decides where a request goes; <a href="https://www.jonathanmuk.com/insights/rate-limiting">rate limiting</a> decides whether it should be served at all, and production systems need both.</p>
<h2 id="monitoring-metrics"><a href="#monitoring-metrics">Monitoring Metrics</a></h2>
<p>A load balancer is also a great observation point for your system&#x27;s health, since all traffic flows through it.</p>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 13rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Traffic Metrics</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Tell you about volume: how much is hitting your system. RPS (requests per second) is the primary heartbeat. A sudden spike could mean a viral moment or a DDoS attack.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Performance Metrics</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>The most actionable for user experience. Average response time is misleading on its own. P95 and P99 tell you what the slowest 5% and 1% of users experience. If your average is 50ms but P99 is 5 seconds, something is deeply wrong for a non-trivial slice of users.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Health Metrics</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>The load balancer&#x27;s own view of your servers. Load balancers send periodic health check requests (e.g. GET /health) and mark servers as unhealthy if they fail. Once unhealthy, traffic is stopped to that server automatically.</p></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Error Metrics</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>The distinction between 4xx and 5xx is critical. 4xx (client errors) are usually application-level issues. 5xx (server errors) are infrastructure problems. A rising 5xx rate is a red alert.</p></div></li></ul>
<blockquote>
<p>Load balancers let you scale horizontally and give you redundancy. If one server crashes, traffic is automatically rerouted to the healthy ones. That is the foundation of any production system at scale.</p>
</blockquote>]]></content:encoded>
            <author>mukjonas256@gmail.com (Jonathan Mukhobe)</author>
            <category domain="https://www.jonathanmuk.com/insights?tag=Load%20Balancing">Load Balancing</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=System%20Design">System Design</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Scalability">Scalability</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=NGINX">NGINX</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Cloudflare">Cloudflare</category>
            <enclosure length="106850" type="image/jpeg" url="https://www.jonathanmuk.com/og/insights/load-balancing.jpg"/>
        </item>
        <item>
            <title><![CDATA[Rate Limiting: The API Guard You Cannot Ship Without]]></title>
            <link>https://www.jonathanmuk.com/insights/rate-limiting</link>
            <guid isPermaLink="false">https://www.jonathanmuk.com/insights/rate-limiting</guid>
            <pubDate>Thu, 22 May 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[How rate limiting and throttling protect APIs: token bucket against leaky bucket, where to enforce each, and why Redis is needed once you run several nodes.]]></description>
            <content:encoded><![CDATA[<h2 id="rate-limiting-and-throttling-on-api-endpoints"><a href="#rate-limiting-and-throttling-on-api-endpoints">Rate Limiting and Throttling on API Endpoints</a></h2>
<p>An unprotected API endpoint is an open invitation to abuse, from scrapers and bots to cascading backend overload. Understanding how rate limiting and throttling work is non-negotiable for any backend engineer shipping production systems.</p>
<h2 id="rate-limiting-vs-throttling"><a href="#rate-limiting-vs-throttling">Rate Limiting vs Throttling</a></h2>
<p>Think of an API as a highway. Rate limiting is a hard barrier, a toll gate that only lets through a fixed number of cars per hour, then turns the rest away entirely. Throttling is a speed bump: it does not block anyone, it just slows the convoy down to a manageable pace.</p>
<p>Both protect your system, but they are different philosophies. Rate limiting enforces a quota; throttling manages flow.</p>
<ul class="not-prose my-8 grid list-none gap-3 p-0" style="grid-template-columns:repeat(auto-fit, minmax(min(100%, 16rem), 1fr))"><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Rate Limiting</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Setting a cap on the number of requests a client or user can make within a specific time window (e.g. 1,000 requests per minute per user). If a client tries to exceed this limit, they get blocked or receive an error response, typically a <strong>429 Too Many Requests</strong> status code with a Retry-After header. This ensures no single source overwhelms the system.
You can apply limits at different scopes:</p><ul>
<li><strong>Per-user:</strong> each user gets 1,000 requests per minute, preventing one bad actor from ruining others.</li>
<li><strong>Per-application:</strong> the entire service handles 50,000 requests per minute, protecting the whole system.</li>
<li><strong>Per-endpoint:</strong> stricter limits on expensive endpoints (e.g. /generate-report gets 10/min, /health gets unlimited).</li>
<li><strong>Per-tier:</strong> free users get 100/hour, paid users get 10,000/hour. This is the classic SaaS model.
Rate limiting is ideal when you want to enforce strict usage policies. It is straightforward: you hit the cap, you get denied.</li>
</ul></div></li><li class="rounded-md border border-line bg-surface-1 p-5"><p class="!mt-0 !mb-2 text-h5 font-semibold text-heading [overflow-wrap:anywhere]">Throttling</p><div class="text-small leading-relaxed text-body [overflow-wrap:anywhere] [&amp;&gt;*]:my-0 [&amp;&gt;*+*]:mt-2 [&amp;_code]:text-[0.9em]"><p>Throttling is about slowing down the rate of requests rather than outright rejecting them. It is useful when you want to maintain service availability but need to manage load more gracefully. If an API is handling more requests than usual, you can queue requests and process them at a steady rate instead of rejecting them. The user might see increased latency, but they are not blocked outright.
Throttling does not say &quot;you cannot,&quot; it says &quot;wait your turn.&quot; Your client experiences latency but eventually gets a response.</p><ul>
<li><strong>Dynamic throttling:</strong> adjusts levels based on real-time metrics like CPU load, memory usage, or queue depth. When your server is at 90% CPU, you slow the drain on the queue. When it drops to 40%, you speed back up.</li>
<li><strong>Adaptive throttling:</strong> takes this further by using heuristics or ML to predict spikes before they hit. If every Monday at 9am traffic triples, an adaptive system starts throttling at 8:55am, before the wave arrives.</li>
</ul></div></li></ul>
<p>Often systems implement both: first try throttling, then move on to rate limiting if capacity is threatened.</p>
<h2 id="the-two-core-algorithms"><a href="#the-two-core-algorithms">The Two Core Algorithms</a></h2>
<h3 id="token-bucket-rate-limiting"><a href="#token-bucket-rate-limiting">Token Bucket (Rate Limiting)</a></h3>
<p>This is the most widely used algorithm for rate limiting. Imagine a bucket that holds tokens, where each token represents permission to make one request. Tokens are added to the bucket at a fixed rate, and if the bucket is full, extra tokens are discarded. When a request comes in, the program checks the bucket: if a token is available, the request is allowed and a token is removed. If the bucket is empty, the request is denied.</p>
<p>The token bucket is great for rate limiting because it naturally handles bursty traffic. A user who makes no requests for 30 seconds builds up a reserve of tokens. When they need to fire a batch of requests, they can. This is more fair and realistic than a rigid counter.</p>
<p>In practice (e.g. in Redis), you store two values per user: <code class="inline-code">tokens_remaining</code> and <code class="inline-code">last_refill_timestamp</code>. On each request, you calculate how many tokens to add since the last refill, clamp to max capacity, then check if there is a token to consume.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><img src="https://www.jonathanmuk.com/figures/img4.svg" alt="Token Bucket algorithm diagram" width="1200" height="800" loading="lazy" decoding="async" class="h-auto w-full"/></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Token Bucket algorithm diagram</figcaption></figure>
<h3 id="leaky-bucket-throttling"><a href="#leaky-bucket-throttling">Leaky Bucket (Throttling)</a></h3>
<p>The leaky bucket solves a different problem: smoothing out a burst rather than capping it. Imagine a bucket that leaks water at a constant rate. When a request comes in, it is treated like adding water to the bucket. If the bucket is not full, the request is allowed and a bit of water is added. If the bucket is full, the request is delayed or rejected. The bucket ensures that requests are processed at a smooth constant pace, even if lots arrive at once.</p>
<p>The leaky bucket is ideal when you are calling a downstream service that can only handle a steady rate: a payment processor, a third-party API with strict SLAs, or a legacy backend. You absorb the burst, then deliver requests calmly. Your client waits a bit, but the downstream service stays healthy.</p>
<figure class="my-10"><div class="overflow-hidden rounded-md border border-line bg-surface-1"><img src="https://www.jonathanmuk.com/figures/img5.svg" alt="Leaky Bucket algorithm diagram" width="1200" height="800" loading="lazy" decoding="async" class="h-auto w-full"/></div><figcaption class="mt-3 font-mono text-caption leading-relaxed text-muted">Leaky Bucket algorithm diagram</figcaption></figure>
<h2 id="the-full-implementation-picture"><a href="#the-full-implementation-picture">The Full Implementation Picture</a></h2>
<p>Systems often layer both approaches: throttle first to queue the burst, then rate limit to reject genuinely excessive clients. Here is what that looks like end-to-end:</p>
<h3 id="cloudflare-domain-layer"><a href="#cloudflare-domain-layer">Cloudflare (Domain Layer)</a></h3>
<p>This is the first and cheapest line of defense. You can block entire IP ranges before the request even touches your servers. Configure rate limiting rules under Security in the Cloudflare dashboard: for example, &quot;block IPs that make more than 500 requests per minute to /api/*.&quot; This costs nothing in compute and handles volumetric attacks and bot traffic before they reach your backend. Enabling Bot Fight Mode handles a huge class of abusive traffic automatically.</p>
<h3 id="vercel-frontendedge-layer"><a href="#vercel-frontendedge-layer">Vercel (Frontend/Edge Layer)</a></h3>
<p>Vercel&#x27;s Edge Middleware runs before your React app and is a great place to rate limit session-based requests or SSR routes. You can write middleware that checks a counter in an edge KV store. For most React apps, you will rely on the backend for heavy rate limiting.</p>
<h3 id="fastapi-main-enforcement-layer"><a href="#fastapi-main-enforcement-layer">FastAPI (Main Enforcement Layer)</a></h3>
<p>Use <code class="inline-code">slowapi</code>, the FastAPI equivalent of Express Rate Limit. You can apply per-IP limits globally or set stricter per-endpoint limits for expensive operations like report generation. For user-based limits (not just IP), swap the key function to extract the user ID from your JWT.</p>
<h3 id="django-rest-framework"><a href="#django-rest-framework">Django REST Framework</a></h3>
<p>DRF has built-in throttling via <code class="inline-code">AnonRateThrottle</code> and <code class="inline-code">UserRateThrottle</code>. You configure default rates in settings.py or create per-view throttle classes for endpoint-specific limits. DRF uses a leaky bucket-style algorithm internally, which is great for throttling.</p>
<h3 id="redis-shared-state-across-instances"><a href="#redis-shared-state-across-instances">Redis (Shared State Across Instances)</a></h3>
<p>The moment you scale to more than one instance, in-memory rate limiting breaks: each instance has its own counter and they do not talk to each other. Redis solves this. Add a Redis instance and point both slowapi and DRF at it. Both libraries support Redis backends directly, keeping your counters consistent across every running instance.</p>
<p>Rate limiting sits alongside <a href="https://www.jonathanmuk.com/insights/load-balancing">load balancing</a>: one caps what a client may ask for, the other spreads what gets through.</p>
<h2 id="monitor-and-tune"><a href="#monitor-and-tune">Monitor and Tune</a></h2>
<p>Set up alerts on your 429 rate. If it spikes, you might be too aggressive (blocking legitimate users) or too loose (absorbing abuse). Track these metrics:</p>
<ul>
<li>4xx error rate, filtered to 429s.</li>
<li>P95/P99 latency (throttling shows up here before errors do).</li>
<li>Queue depth if you are using any async processing.</li>
<li>Per-user request rates to spot abusers.</li>
</ul>
<p>For metrics, Railway&#x27;s built-in observability dashboard is a start, but for production you will want to push to Datadog or use a lightweight option like Prometheus and Grafana.</p>
<h2 id="quick-decision-guide"><a href="#quick-decision-guide">Quick Decision Guide</a></h2>
<div class="table-wrap not-prose my-8"><table class="w-full text-small"><thead><tr><th scope="col" class="border-b border-line-strong px-3 py-2 text-left font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Situation</th><th scope="col" class="border-b border-line-strong px-3 py-2 text-left font-mono text-eyebrow font-medium tracking-[0.02em] text-muted">Use</th></tr></thead><tbody><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Free vs paid tier quotas</th><td class="border-b border-line px-3 py-2.5 text-body">Rate limiting (token bucket)</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Protecting a slow downstream API</th><td class="border-b border-line px-3 py-2.5 text-body">Throttling (leaky bucket)</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">DDoS / bot traffic</th><td class="border-b border-line px-3 py-2.5 text-body">Cloudflare rate limiting (IP-level)</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Single-instance app</th><td class="border-b border-line px-3 py-2.5 text-body">In-memory (slowapi default)</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Multi-instance on Railway</th><td class="border-b border-line px-3 py-2.5 text-body">Redis-backed rate limiter</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Expensive endpoints (AI, reports)</th><td class="border-b border-line px-3 py-2.5 text-body">Per-endpoint stricter limits</td></tr><tr class="even:bg-page-subtle"><th scope="row" class="border-b border-line px-3 py-2.5 text-left font-medium text-heading">Prefer user experience over strictness</th><td class="border-b border-line px-3 py-2.5 text-body">Throttle first, rate limit as fallback</td></tr></tbody></table></div>
<blockquote>
<p>The general rule: Cloudflare handles the perimeter, FastAPI or Django handles business logic limits, and Redis keeps state consistent when you scale.</p>
</blockquote>]]></content:encoded>
            <author>mukjonas256@gmail.com (Jonathan Mukhobe)</author>
            <category domain="https://www.jonathanmuk.com/insights?tag=Rate%20Limiting">Rate Limiting</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Throttling">Throttling</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=API%20Design">API Design</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Backend%20Engineering">Backend Engineering</category>
            <category domain="https://www.jonathanmuk.com/insights?tag=Redis">Redis</category>
            <enclosure length="132635" type="image/jpeg" url="https://www.jonathanmuk.com/og/insights/rate-limiting.jpg"/>
        </item>
    </channel>
</rss>