MODEL CONTEXT PROTOCOL AND A VISION OF HISTORICAL RESEARCH AND COLLECTIONS ACCESS
The following post includes my remarks (lightly edited) that I delivered for a 2026 American Historical Association conference panel, “AI-Mediated Historiography: Models, Misinformation, and Methodology,” with fellow panelists Jana Dunz-Keck, Todd Presner, and Fabian Offert.
During the seventeenth century, a market for buying and selling exotic tulips emerged in the Dutch town of Haarlem as it did throughout much of the Netherlands. A rising middle class of merchants built walled gardens within and beyond the town’s walls to collect, study, and admire rare varieties of tulips that proved difficult to cultivate. At the same time, artists such as Ambrosius Bosschaert the Elder and his brother-in-law Balthasar van der Ast captured the tulips’ mystery and beauty in enchanting still lifes, in many cases pairing delicate tulips with shells, as seen above in one of van der Ast’s paintings. Both shells and tulips were equally sought-after for their uniqueness and scientific interest.
Now, imagine you wanted to develop a lesson plan, write a college essay, or conduct original research that analyzes the sociocultural connections between the Dutch Golden Age of art and the flashpoint moment in economic history known as Tulip Mania that erupted in the mid-1630s. Like it or not, more people are turning to large language models as a go-to source not just for inspiration, but for information and research assistance.
I’d like to recount a recent exchange with a large language model that may feel all-too familiar. I asked ChatGPT 5.2 Pro about the historical significance of Tulip Mania, the brief moment popularly remembered for when the price of some rare tulips were exchanged for as high as the price of a home. ChatGPT delivered a short synopsis that seemed helpful enough. In a series of bullet points, it described Tulip Mania as “an early example of a speculative bubble” and acknowledged that recent scholarship has mostly debunked the theory that Tulip Mania constituted a nationwide financial collapse.
At the top of the response, was a set of uncited images of artwork depicting tulips supposedly from the period.
A follow-up question asking for the source information for the works yielded an admission that ChatGPT did not know anything about the images it displayed and that they were “not guaranteed to be contemporaneous depictions of the events.”
Not even an image’s appearance in Wikipedia’s page on tulip mania, which identified the plant above from a 1637 Dutch catalogue as “the Viceroy”, could lead ChatGPT toward providing a citation and confidently placing it squarely within the historical moment. After ChatGPT volunteered to create an annotated list of images, it still could not accurately identify any of the works that it collated.
This exchange encapsulates both the promises and perils of using LLMs for historical research. In one sense, ChatGPT’s initial synopsis of Tulip Mania did not veer wildly into misinformation or hallucinations and seemed modestly aware of current historiography on the subject driven by historian Anne Goldgar’s scholarship. However, as soon as I began to query the chatbot for further source information, its reliability collapsed. LLMs may often identify relevant secondary scholarship if prompted, but they clearly struggle with connecting their synthetic summaries with primary source material, which has considerable consequences whether you are a teacher developing a lesson plan, a student writing a report, or researcher conducting a preliminary search.
Today, I would like to suggest an emerging standard called Model Context Protocol, or MCP, that may fundamentally alter the AI history landscape. I will explain how MCP works and present a brief hypothetical example developing a lesson plan for a middle school art class. MCP offers the potential for improving LLM reliability by facilitating direct connections with collections data from libraries, archives, and museums, which, over time, could open novel historical research methods. Rest assured, the vision I’m proposing today, if executed properly, will not be the demise of our livelihoods; rather it will elevate the need for human historians, archivists, librarians, and educators.
First, let me explain (in a simplified way) how MCP works. Think of MCP as the protocol that establishes a gateway between a database and an LLM. You ask a question in the form of a chatbot text prompt, and the LLM retrieves the relevant data directly from the host database.
MCP has three layers. At one end is your AI application, typically a LLM such as Claude or ChatGPT. At the other end is the data source. In-between is the MCP server, which acts as a translator, exposing specific tools the LLM can use to query the data. When you ask a question in the form of a prompt, the chatbot interprets what you are seeking and decides which tools to employ. The server fetches the data, and the chatbot then synthesizes it into a response.
MCP was first developed by Anthropic in late 2024. A year later, they donated MCP as open source to the Linux Foundation which announced the formation of the Agentic AI Foundation. According to the site, the Foundation’s mission is to “provide a neutral, open foundation to ensure agentic AI evolves transparently and collaboratively.” All the major frontier labs, including Google, OpenAI, and Microsoft have joined the collective, further establishing MCP as a universal standard protocol, which is good news for the heritage community because it does not lock them into any single proprietary system.
So, what is the significance of MCP for historians? In its most ideal state, MCP can enable a library, archive, or museum to connect its collections data to an LLM such as Claude, Gemini, or an assortment of open-source models. This represents the first step of the dream scenario in which a researcher can use a computer to “chat” with a historical collection regardless of its size and scope.
Here’s how the MCP exchange I explained a moment ago works in the context of historical collections. Imagine typing a query requesting information pertaining to the collection and the model accesses a designated institution’s collection data, whether in the form of descriptive metadata, catalog records, finding aids, full text, or the digital objects themselves. The advantage is that unlike traditional search, which relies upon identifying the exact terms that you hope will yield optimal results, the LLM allows for more open-ended or ambiguous queries. An MCP-directed exchange can translate your historical query, however vague, and open multiple entry points into the data. Like typical LLM chats, a user can engage with the LLM in a more conversational context, asking follow-up questions and probing for additional information that may otherwise not surface as readily in a typical keyword search approach.
Among over 15,000 MCP servers (and counting) currently available, there are only a handful of examples in the library, archives, and museum space, but Ruud Huijts developed an MCP server for the Rijksmuseum in the Netherlands that might serve as a model for other heritage institutions. The following is one of my early encounters with the MCP server seeking examples of works depicting tulips. (I should note that for this early exchange, I had not settled on researching Tulip Mania; more on that in a moment).
The first thing to point out is that before I could conduct a query, I needed to connect the MCP server to the local chatbot on my desktop, in this case the latest version of Claude. Thankfully, the developer provides useful step-by-step instructions on his Github page on how to paste a batch of code into Claude’s backend files. The Rijksmuseum provides guidance on how to obtain an automatically generated free API key, which was necessary to gain permission to access the database.
One advantage was that whenever I encountered minor hiccups along the way, Claude was extremely helpful in identifying missteps given its strength in coding. You can see here that a successful connection in Claude is represented as a “connector.”
Once I had successfully installed the MCP server on my computer, I faced the daunting task of asking a question, not unlike approaching an empty search bar.
For all its strengths, the Rijksmuseum does not provide a comprehensive guide or catalogue that outlines the scope of its collection. At various points on its site, it suggests the collection contains over 800,000 objects, but unless you are an art historian or frequenter of the museum with extensive knowledge of its holdings, you may have difficulty gauging what kinds of queries you may be able to pose. Beyond Rembrandt, I had no idea which artists, periods, or types of art would yield results.
On a whim, I typed a simple prompt requesting high-resolution images of works containing tulips. It identified 14 works and displayed four in my web browser, beginning with three still lifes from the Dutch Golden Age and a nineteenth-century impressionistic painting of a laborer working in a tulip field.
It is important to point out that a standard search for “tulip” on the museum’s main site indicated there are 616 artworks, so right away we can see that there is a question of whether the first 14 works identified by the chatbot were completely random or curated according to some unknown set of criteria.
Whether the output was gently nudging me in a particular direction or by sheer coincidence, the chronological and stylistic proximity among the three still lifes compelled me to engage with them in more depth.
Next, I asked Claude to develop a lesson plan for a middle school art class based on the three paintings. This is where the real jaw-dropping moment occurred.
Without my specifying any historical connection with Tulip Mania, in moments Claude drafted a ten-page, four-day plan that frames the paintings within the historical moment and encourages close examination of the works.
The opening question prompts students to consider how much they would pay for a tulip seen in the paintings, followed by a 15-minute “mini-lecture” that discusses Dutch trade, the rise of new merchant class, and the speculation “craze” surrounding exotic tulips. Subsequent sections of the lessons briefly situated the three artists, Hans Bollongier, Ambrosius Bosschaert, and Balthasar van der Ast within a broad historical context, before walking students through a step-by-step art analysis covering the works’ composition, use of light and shadow, and symbolism, including the appearance of shells and insects. The multi-day activity culminates with students thinking about present-day consumerism and instructs them to create their own still life using symbols that represent their tastes. What I found most striking was Claude’s brief rationale for how it structured the lesson plan. Taking into consideration the age level of the students, it commented:
“The approach makes 400-year-old paintings relevant by connecting to things middle schoolers care about - what’s valuable, what’s trendy, what represents them - while teaching real art analysis skills and historical context!”
This entire exchange occurred in under five minutes.
So why aren’t historians, librarians, archivists, curators, and teachers all doomed? The answer lies in how a MCP server accesses data. When you peek under the hood, you will see that the server does not enable completely open-ended access to the collection. According to the Rijksmuseum MCP Github page, there are currently six possible types of interactions called “tools.” You can conduct a standard search for basic catalog information about an artist or work such as a work’s title or creation date. You can go deeper into a work to retrieve “comprehensive information” about its “physical properties,” “historical context,” “exhibition history,” and other provenance information. Additional tools allow you to create custom collections, view images either in a standard web browser or at high-resolution with zoom capabilities, or develop a timeline that tracks artistic periods and artists’ stylistic development.
You can confirm these functionalities, and even customize which ones are accessible, directly within the Claude Desktop app.
Each of these possible routes into the collection was a deliberate choice made by the developer and determined by a combination of the museum’s data policy and the quality of its metadata.
In other words, MCP enables--even empowers--an institution to maintain intellectual control over its data. Users no longer are reliant on trusting chatbot responses prone to hallucinations and distortions. If LLMs represent the unpredictability of a wild, untamed landscape, MCP creates a walled garden in which your prompts and the information retrieved remain within the boundaries dictated by the scope of the collection and a set of instructions determining how to access the data.
The key here is that the MCP exchange significantly reduces the LLM’s reliance on probabilistic outputs. With a typical prompt, it is difficult to gauge which sources contributed to the output, whereas using MCP, the chatbot prioritizes the inscribed data above the data upon which it was trained. Unlike the ChatGPT response I showed earlier, an image retrieved from the museum’s database contains an accurate catalog record.
Librarians, archivists, curators and other stewards of collections therefore are more, not less, critical in securing collection access via MCP. The quality of the outputs that you can generate with a prompt are dependent on the quality of the metadata. The Rikjmuseum’s MCP server only works as well as it does because of the meticulous curation of the museum’s catalogue records built upon Linked Open Data. In other words, developing an MCP server is a variation of arranging and describing a collection for enhanced access. It provides the scaffolding to think through issues of intellectual control within generative AI, which requires thoughtful consideration for how to merge standards, humanistic analysis, and practices to optimize results for research, education, and public programming.
The robustness of the Rijksmuseum’s MCP server is reflected by the care with which the museum has rolled out access to its data. Since December 2024, it has developed API interoperability with three distinct catalogs labeled the Rijksmuseum Collection, Research Library, and Object Documentation. There are plans this year to integrate its Research Data collection and in 2027 its Institutional Archive.
As for educators, as tempting as it may be to create a lesson plan in minutes, we are not a point where we can and should hand over teaching history to AI chatbots. In my scenario, the Claude-generated lesson plan appears to cover topics at a level appropriate for a middle school audience, such as the rise of the Dutch merchant class and tulips as a status symbol. But look a little more closely, and the lesson plan reinforces the mostly debunked theory of widespread economic collapse.
When I pointed this out to Claude Sonnet 4.5 Pro, it acknowledged it may have over-generalized and proceeded to conduct a seven-minute research query that supposedly consulted over 300 sources. Its revised summary was far more accurate and detailed, but what was even more interesting was the explanation it provided for why it leaned heavily on Tulip Mania myth more than fact:
“The persistence of tulip mania mythology in AI systems reflects the sources these models were trained on. Economics textbooks, popular finance books, and general reference materials overwhelmingly repeat [Charles] Mackay’s [1841] narrative. Claude—like other AI systems—may perpetuate these myths unless specifically prompted with revisionist scholarship.”
In other words, the predominance of Internet texts used for the model’s training, which compare cryptocurrency and other speculative bubbles to the most exaggerated or disproven claims about Tulip Mania, risks outweighing the far fewer pieces of quality scholarship.
This suggests that as hard-pressed for time teachers may be, they still must conduct their due diligence to ensure curriculum needs are met while avoiding perpetuating outdated arguments. Even a few simple follow-up questions can weed out misleading or outdated claims.
These are certainly still early days, and many questions and years’ worth of dedicated work remain. Foremost, while an MCP server theoretically narrows outputs to information stored in the database, it is not entirely clear the extent to which hallucinations can still creep in. As far as I can tell, for example, Claude took the initiative in associating the selected still life paintings with Tulip Mania. A keyword search on the museum’s collection search interface yielded no results about the historical moment, so it’s likely the information came from Claude and not the MCP server. LLMs are synthesizing model data with data retrieved from the museum in ways that may be difficult to parse. In this instance, while the accurate art-related information was retrieved from the museum catalog, Claude took the liberty of wrapping that information with its own layer of historical context. How do we distinguish the line between model versus database-provided information?
Second, there is the issue of maintaining quality metadata. Very few institutions have the resources to steadily build an open access collection at the level of the Rijksmuseum, which claims to have been working on improving digital access for over 12 years. Most institutions maintain records that contain only basic metadata, which raises the question of how differing levels and approaches to data will impact the utility of a MCP server. On the reverse end of the spectrum, how well do MCP servers hold up when accessing full-text databases? Is the context window of current LLMs robust enough to probe entire digital libraries?
Third, as I mentioned at the opening, whatever systems are developed, they must facilitate reliable citation. When historical objects are retrieved, will researchers be able to document source locations? To what extent should the MCP interaction itself be documented?
Finally, we must not overlook the environmental and financial costs that accompany any sustained use of AI-enabled technologies. As MCP and other tools and standards undergo continued development, they invite extended use sessions that will likely grow in complexity, duration, and cost that can place strains on smaller institutions with limited capacity.
I see these questions as not necessarily obstacles, but as collaborative development opportunities for historians and heritage practitioners.
I would like to leave you with a slightly longer-term vision of where MCPs might take historical work. Imagine a time in the hopefully not-too-distant future where you can conduct research across disparate collections. As more institutions develop MCP access to their data, users will be able to line up multiple MCP servers within their chatbot, which would open potential research methods that traverse collections.
For example, you may be interested in comparing governmental response to the 1918 influenza outbreak at the federal, state, and local levels. You might line up access to the National Archives, Smithsonian, Library of Congress, various state archives, and municipal archives. Traversing collections across institutions simultaneously, the LLM could synthesize information that could point you to hidden collections for further investigation and comparative frameworks that may have taken months and multiple trips to archives to establish. Far from replacing historians, AI models connected to these walled gardens of carefully curated collections could unlock creativity and spur novel arguments grounded in historical evidence.












