Preserving Trust in AI Knowledge Ecosystems
Keynote remarks delivered at the 2026 AI Summit hosted by the University of Wyoming Libraries
On May 29, 2026, I delivered the keynote remarks for the University of Wyoming Libraries’ 2026 AI Summit. The day-long event had a wonderful turnout of regional librarians, archivists, teachers, and technologists. I encourage you all to check out the full program! Also, some of my examples reference material from previous Encoding the Past posts. You can find a fuller discussion of the Rijksmuseum case study here, and Parts 1 and 2 of my Chronicling America Jenny Lind Tour case study (Part 3 forthcoming!).
Table of Contents
Before I begin, I want to take a moment to thank the people and institutions who made this summit possible. I’m especially grateful to Sara Davis, Abby Beaver, Paula Martin, and Lucy Carter for their insights and hospitality.
It is also worth pausing to recognize the collaboration represented here among the Wyoming State Archives, the Wyoming State Library, and the University of Wyoming Libraries. That kind of partnership across state heritage institutions is not something that happens every day, and it feels especially appropriate for a conversation about trust, stewardship, and the future of knowledge ecosystems.
I’m honored to be part of that conversation with you today.

A Wild West for Authentication
Imagine you're presented with this famous artwork, Lion on the Watch by Jean-Léon Gérôme, and asked to establish its full provenance. To trace the work's history, beginning from when it was painted to when it landed in the Cleveland Museum of Art’s collection, you might research artist correspondence, diaries, or earlier sketches.
As it changed hands, you'd consult exhibition catalogs, sales ledgers, and estate records. You might even take multispectral images or small samples for material analysis to reveal whether the work was altered by the artist or a prior conservation treatment. The Getty Research Institute reconstructed a full history of the painting, but in cases where ownership is unaccounted for, uncertainty about authenticity follows. The more complete the portrait you can compile of a work’s provenance, the more you can trust that it is what it purports to be.

Now imagine trying to account for the provenance of a digital object today. It’s not a stretch to conclude that this image of penguins in a desert was fabricated, but we don’t know how or by whom. With sophisticated generative AI at everyone’s disposal, determining an object’s origin, or whether it has been altered, is genuinely difficult. Someone can be erased from an image or have their voice synthetically replaced. Digital content crosses platforms in seconds, each hand-off a potential site of modification. AI has introduced a Wild West for authentication that has sown doubt in the foundations of what we hold to be true.
Late-breaking news flashed across my inbox as I was finishing these remarks, which goes to show that you should never write a talk about AI more than a week out! Within minutes of each other last week, Google and OpenAI announced new commitments to content provenance using two complementary approaches: one an open standard, the other an invisible watermarking technology. The first, developed by the Coalition for Content Provenance and Authenticity, or C2PA, attaches cryptographically signed metadata known as Content Credentials to a digital object such as an image. The second, SynthID, created by Google, embeds an invisible watermarking layer within the object that can be detected by tools designed to identify a company’s synthetic content. Upload an image to their beta tool, and they will determine whether they believe the image has been generated by their AI model.
This short video by Google demonstrates verification in action.
The Central Currency of Trust
The official adoption of these two approaches, C2PA and SynthID, marks a turning point in the industry, and potentially more broadly in society, in prioritizing digital provenance. Rather than guessing its origin based on makeshift detection tools, the big tech companies are moving toward a shared provenance ecosystem by which metadata about a digital object’s creation and subsequent alterations can travel across platforms. Before long, we will become accustomed to seeing a label, such as C2PA’s content credentials, that explains the history of a digital object, including whether an AI tool placed a penguin in the Saharan desert.
OpenAI heralded the announcement as a “multi-layered, ecosystem-driven model to building trust online.” However, this development represents merely the first step toward a long journey to secure trust in digital content via authenticity and provenance, and it is a journey we all must embark upon together. Trust is the central currency by which libraries, archives, and museums operate. If AI erodes trust, as we are witnessing within so many other sectors of our society, we lose a critical connection to our past and our capacity for planning the future.
Today, I would like to consider how content authenticity and provenance, or what I will refer to as CAP, sustain trust through the framework of knowledge ecosystems. Knowledge ecosystems refer to the interconnected network of people, institutions, tools, standards, collections, and interpretive practices through which information becomes trusted knowledge. At the heart of the ecosystem represented in the blue circle on the left are libraries, archives, and museums, or LAMs, what I refer to as the trust keepers of knowledge; in the purple circle on the left are the users, whom I call the knowledge producers.
But here is what makes AI’s disruption of that ecosystem distinctive — and, I would argue, underappreciated. The more AI interacts with heritage data, the more it injects additional layers of provenance into the equation. LLMs and other types of AI models carry their own knowledge sets based on their pretraining. Those training datasets possess distinct provenance, which ought to raise questions about how they were selectively compiled by considering their geographic, cultural, and temporal coverage. They do not encapsulate the entirety of human knowledge. Instead, like any other dataset, they were curated, mostly from the Internet and readily available digitized works, representing a snapshot in time. In other words, we must account for the interactions of two separate sets of data provenance: heritage collections and the AI models that interact with them. Along the way, the models introduce several points where their pretrained knowledge can intervene and endanger our confidence in the results. I call this phenomenon provenance laundering, and I will return to it later in my talk.
It is easy to get swept up in the latest tools or model releases, but instead, what I would like to offer today is a framework to sustain trust in knowledge ecosystems. I will begin with a survey of where we are as a field and how we got here. Knowledge ecosystems work best when they are grounded in three core values: authenticity, provenance, and community, yet responses to working with AI focus heavily on community while overlooking the other two values. A brief series of vignettes will demonstrate how AI’s disruption to authenticity and provenance over the last decade contribute to the fraying integrity of knowledge ecosystems. The examples will likely feel familiar at first, but as we move to the present, they may start to feel increasingly strange. The next section will discuss CAP from the perspective of LAMs as the trust keepers of knowledge, drawing on work I conducted with my colleague Kate Murray at the Library of Congress, which culminated in an open access white paper earlier this year. In the third section, I will consider CAP from the point of view of knowledge producers such as researchers, students, educators, and community stakeholders, before a brief conclusion about where all this may be heading.
If there is one idea I would like to leave you with for today's summit, it is that we must not lose sight of authenticity and provenance, which comprise the core values that should govern our mission. Sustaining trust within knowledge ecosystems will require a renewed commitment among preservation practitioners and users to cooperate in building systems and standards grounded in these core values.
From Deepfakes to Agents: How We Got Here
Before we turn to each of the two halves of the knowledge ecosystem, LAMs and users, I wanted to sketch roughly current responses to AI and how we arrived at this moment. The heritage field has made remarkable strides establishing community as the central value for addressing AI’s risks and opportunities, which is understandable and laudable given its history of expanding access. I have observed firsthand remarkable research and development projects advancing Indigenous ways of knowing, elevating minoritized voices, and expanding accessibility that have transformed how we steward knowledge.
That community-centered framework has shaped the field’s early response to AI. Across the country, professional associations, universities, and LAMs have rushed to publish AI guides and competency frameworks that foreground ethics including algorithmic bias, environmental costs, and accessibility gaps. These efforts reflect genuine urgency and good professional instincts. But as a rule, they move quickly past two values that have anchored our work since long before AI arrived.

From seventeenth-century diplomatics to digital forensics, authenticity and provenance have been the bedrock of heritage work through every technological upheaval. A researcher requesting an archival document, whether a medieval manuscript or the files from a personal hard drive, wouldn’t consult an archivist to determine whether the document is accurate but rather would expect assurance that the document is an original.
In the blink of an eye, we have lost our bearings on what is authentic and how we can prove it. I read just the other day in an op-ed that when a parent showed her child a photo of a beautiful landscape taken during a vacation, the child instinctually responded that it was fake. It is astounding how quickly we have come to assume that content is AI-generated.
So how did we arrive at this moment of extreme skepticism? Let’s begin with the quaint period of the late 2010s and early 2020s, when our top concerns were occupied by AI-generated media content posing as authentic, commonly known as deepfakes. Deepfakes of foreign leaders, the Pope, wars, and numerous other sensitive topics have flooded our social media feeds, news outlets, and academic discourse, introducing concern about how to mediate a media-soaked environment filled with fictitious content. Some were transparently experimental, such as the example here of a deepfake Richard Nixon reading actual prepared remarks had NASA’S Apollo 11 team not returned.
The arrival of generative AI in 2022 and the public launch of ChatGPT only accelerated the trajectory toward content falsely purporting to be authentic. Overnight, legal rulings, government policies, political speeches, student assignments, and scholarly publications all exhibited tell-tale signals of poor use of large language models with invented sources, wrong information, shoddy analysis, and unoriginal arguments.
More positively, LAMs have experimented with using AI to preserve and provide access to digitized and born-digital collections. When records are numbered in the millions, the question becomes whether AI can automate processing. For example, recent advances have elevated the accuracy of transcribing illegible handwritten documents, which previously depended upon professional or crowdsourced input page by page. I suspect we’re not far off when the tedious effort to compose a finding aid may be handled in hours versus days.
Flash forward to 2026 and the heritage landscape has only become stranger. Suddenly, users with advanced programming experience or those, such as myself, wholly reliant on coding agents like Claude Code, have experimented with deploying the latest large language models to access LAMs’ collections using acronym-heavy approaches such as Retrieval-Augmented Generation, or RAG, and Model Context Protocol, or MCP. As I will explain further momentarily, within hours, armed with an API key and a vibe-coded set of tools, you can access the entire catalog of the Rijksmuseum or the millions of digitized historic newspaper pages hosted on Chronicling America, all within the comfort of your home anywhere in the world, and often for free or relatively low cost.
At the same time, AI is straining preservation frameworks in ways that we are only beginning to reckon with. Leading tech companies are quietly approaching LAMs, offering to enhance access to collections in exchange for using that content to train frontier AI models. Institutional repositories face the challenge of preserving AI-assisted research outputs such as code bases, prompts, datasets, and experimental results that may constitute the scholarly record of the future. And it won't be long before LAMs must decide whether to acquire collections generated all or in part by an LLM. Before its works were taken down within the last month on Amazon, the fictitious AI author Blake Whiting had over a dozen books about archaeology and ancient history that pirated work by established scholars. Do such works deserve to be preserved, certainly not for their uncited, unattributed research, but perhaps as artifacts of this moment in AI slop?
And things will only get weirder still. The rise of AI agents--autonomous bots designed to conduct tasks with little or no human oversight--will turn the heritage world on its head. In the not-too-distant future, it’s altogether possible that these agents will outnumber human users. While actual humans will continue to trickle in and engage staff at reference desks, hundreds, if not thousands of agents working on behalf of researchers, journalists, or lawyers may be probing your catalogs and collections silently in the background. Already news outlets such as The Economist are reconfiguring their content to accommodate agent-readable versions that can maximize discovery and retain their subscriber base. Whereas we expect content to be presented in a media-rich experience, AI agents work best when content is stripped bare to its structured text.
The sheer pace of development undoubtedly has raised questions for you, whether you are well-steeped in the latest advances in AI or have only heard on occasion that chatbots are eroding our capacity for independent thought, destroying the education system, or alternatively igniting an era of unrestrained growth and prosperity. I’m sure questions have been swirling in your head: Will AI automate away my job? How reliable are these systems, and which one should I use? Will I ever see a living soul visit my institution five years from now? What happens if I do nothing? After all, I have a massive backlog that should keep me busy for the next 20 years!
These are genuinely difficult questions, and I’m not here to prescribe a single AI policy, recommend one model over another, or dictate where AI belongs in your workflows and governance. Every institution is unique. Some will find the terms of working with the private sector acceptable; others will face legal or technical constraints that rule it out. Only you and your colleagues understand fully the context within which you operate.
I can say with some confidence that heritage practitioners such as yourselves are not disappearing, and in fact, I would venture to say that your expertise building and sustaining knowledge ecosystems will prove increasingly central to AI developments. But that confidence depends on all of you taking a proactive stance, and a renewed focus on authenticity and provenance can give us a durable way to reason through choices that otherwise feel overwhelming.
LAMs: The Trust Keepers of Knowledge
Let’s turn to one half of the knowledge ecosystem by considering the central role of LAMs.
As part of a volunteer working group, Kate and I assembled a task force of leaders in the LAMs and technology space, whose names will find in the report. I want to acknowledge their contributions, because this was truly a collaborative effort.
Our starting point was a condition we kept encountering among professionals, particularly at the administrative level, that we came to call praxis paralysis. The institutions I spoke with described feeling caught between competing pressures, with some voices urging quick and decisive adoption of AI or risk being left behind, while others call for caution and deliberation. The prudent approach would be to assess others’ successes and failures, engage in the steady development of resilient standards, and avoid making rash decisions. Unfortunately, there is the shared sense that AI does not allow us to catch our breath and stress-test innovations.
Praxis paralysis, in other words, boiled down to the question: What should my institution do?
Our group acknowledged that LAMs have been engendering trust all along by adapting foundational principles to new technological realities as they arrive. The field has done this before. For example, in the early 2000s, PREMIS established a rigorous data model tracking the full lifecycle of a digital object including what it is, what has acted upon it, and who was responsible. A data dictionary was built and periodically revised. PREMIS is one of several such efforts that, taken together across several decades, gave the field the infrastructure to verify digital content and document its handling, all before AI entered the picture in full force.

What emerged from our review was a set of interrelated challenges to conventional wisdom, beginning with institutional capacity. With each new AI breakthrough, administrators must address anew questions about near- and long-term staffing and skills development. They must assess the value in hiring someone experienced working with AI systems and whether new hires will complement current staff not trained in the latest AI, or whether short-term certifications might provide a temporary stopgap.
Likewise, administrators face daunting questions about upgrading hardware and content management systems. Digital preservation standards already place strains on institutions needing to periodically upgrade servers or LTO systems, conduct file migrations, or apply new metadata schemas to ensure cross-institutional interoperability. What happens when you add AI-enabled tools and platforms that require their own sets of maintenance and upkeep? Approaches to integrating AI into workflows will likely implement a makeshift combination of open-source options, commercial platforms and vendors, or the latest AI coding models. Who will be tasked with ensuring their security? What happens when a single update breaks a tool, or worse, deletes collection data?
Regardless of how AI-centric infrastructure may develop, the second challenge reminds us that they will inevitably run headlong into privacy and ethics risks. Preservation practitioners are discovering how provenance data that records AI-assisted enhancement, reconstruction, or interpretation could inadvertently surface information that conflicts with ethical commitments, donor agreements, or community protocols. In other words, an effort to build in transparency into AI-enabled processes may expose sensitive data. One non-profit organization, WITNESS, developed a Deepfake Rapid Response Force in response to the impact AI-generated or altered content could have on human rights by inciting violence or threatening democracy.
Last year, building on their real-world cases of assessing the authenticity of media content in countries such as Ukraine, Mexico, and India, WITNESS developed the Truly Innovative and Effective AI Detection, or TRIED, Benchmark, a checklist designed to evaluate whether detection tools are held accountable. What you see here are a sample of questions in the TRIED checklist evaluating the extent to which a detection tool could “inadvertently violate human rights, such as privacy or freedom of expression.”

Underlying both capacity and ethics challenges, lies a third, namely the unrelenting pace of change. When Kate and I first released the report back in January, agentic AI was just beginning to gain traction, and it was not entirely clear what the implications would be. Since then, agentic AI has come to dominate discourse and development. The idea is that users will deploy agents with specific capabilities to interact with other agents in data generation, retrieval, and analysis. As agents beget more agents, the pace of change will strain our ability to respond.
Taken altogether, the three challenges paint a picture of institutions needing to assess capacity needs on an ever-shrinking timescale, against a backdrop of ethical, privacy, and security risks that won’t pause for them to catch up. They suggest that, given AI’s pervasive disruption, a traditional response to technological innovation modeled after digital preservation practices may not be sufficient. AI agents are rapidly rewriting the information landscape, necessitating, at times recklessly, entirely new schema, standards, and workflows.
In terms of authenticity and provenance, beyond considering how to secure the provenance of heritage data, we must also account for provenance data that the agents themselves may bring to the table. After all, these agents presumably will represent an individual, organization, or other entity. How do we verify that they are who they claim to be? How can we trace their provenance?
As you can see, you can quickly spin out on wild tangents, but Kate and I wanted to lay groundwork that could be actionable and recognizable to institutions of all sizes. Everyone has a role to play in this transformative journey, from the small historical society or tribal museum to major archives. Community became the bridge between institutional action and the core CAP mission, helping to guide the sustainability of authenticity and provenance by ensuring that changes reflect the needs, risks, and responsibilities of the people LAMs serve.
We outlined four community-centric pillars for consideration that emphasized collaboration. The first pillar calls for coordinated research and development that balances theory and practice to develop sustainable standards, workflows, tools, and platforms. Contrary to what the hype surrounding AI would have us believe, we are still in the early stages of understanding authenticity and provenance needs. For example, international research teams such as InterPARES Trust have yielded considerable understanding of AI-impacted CAP through numerous case studies, analysis of AI tools, and revisions to long-standing CAP concepts.
The second pillar calls for dedicated partnerships, something familiar to the LAMs community. The advantages of institutional collaboration are numerous, including the potential to build economies of scale, resources, and attention. Several recently assembled groups are addressing authenticity and provenance head-on. For example, the volunteer-driven Digital Object Authenticity Working Group is tackling how to authenticate cultural heritage images. Another group, the Trust in Archives Initiative, has designated an Authentication Working Group to “help archives of all sizes evaluate available options and implement the most effective methods of authentication and attestation for their unique contexts.”
The third pillar encourages the LAMs community to leverage their expertise and reputation as trusted institutions to advocate with industry and vendors for rigorous CAP standards and protocols. LAMs are uniquely suited to join existing industry-led groups such as C2PA to develop open and shared specifications. Clearly, big tech recognizes the need for more rigorous provenance standards, but it is unclear the extent to which their vision aligns with the deep knowledge developed by heritage professionals.
The fourth and final pillar calls for a kitchen sink approach to transparent modes of communication. The speed with which AI is evolving inhibits relying upon traditional scholarly outlets as the primary means of sharing information. That is not to say that journal articles, monographs, and anthologies don’t matter, but rather LAMs must use all available modes of instant communication to fill in the gaps. Efforts, some spearheaded by professional societies, should employ blog posts, webinars, Slack channels, and communities of practice based on areas of interest, content, geography, or institution type.
For example, I encourage you all to consider joining AI4LAM, an international collective that meets regularly to share the latest news and insights. AI4LAM holds an annual conference called Fantastic Futures, and this year happens to be in Washington, D.C. in September. If you can’t make it in person, there should be ways of attending virtually.
I have spoken a lot about the need for collaboration from research to tech advocacy. To close this section, I want to clarify exactly what LAMs can bring to the table. All of you understand the importance of preserving context alongside the objects themselves. Context is the human-centered understanding that facilitates lasting knowledge. It cannot be easily automated, and therefore it requires the unique perspectives that all of you bring. Whenever you hear the importance of preserving the human in the loop, I encourage you all to push back, because that implies simply needing someone to oversee AI’s work, to step in when it veers wayward. The value you bring is far richer than ensuring accurate outputs. LAMs occupy the epicenter of unlocking knowledge about human experience by identifying relationships within and among collections. As breakneck developments carry apace, I anticipate that this realization that heritage materials are more than an algorithmic summation of data points will hit the industry hard and LAMs should be prepared to step in as co-equal partners.
Knowledge Producers and The Laundering of Unattributed Knowledge
Let’s turn to the other half of the ecosystem trust equation: users, or knowledge producers of LAMs. I just want to preface this by saying that LAMs practitioners are knowledge producers in their own right. However, they cannot anticipate every scenario of how their information will interact with AI systems, which is where users enter the picture.
As I walk through the issues of authenticity and provenance from the user perspective, I want to highlight a new dimension that is unique to AI-enabled systems, that deserves attention. Besides creating a system of authenticating heritage data, we must also contend with how AI systems intervene in the research process. Using AI models to search and discover content is not a neutral affair. We must also contend with the data, and their provenance, used to train AI models. As we know, pretraining datasets, especially from the frontier labs, tend to be black boxes, so traditional approaches to identifying provenance break down, which leads to what I call provenance laundering. Provenance laundering occurs when knowledge accrued by AI models intermixes with retrieved data, which results in unexpected behaviors at the data retrieval and AI model interpretation stages.
To illustrate this dynamic, I would like to share two case studies I conducted using Model Context Protocol, or MCP, an open standard that connects large language models to databases.
Think of MCP as a gateway between a database and an LLM. You pose a question as a text prompt and the LLM translates that prompt into a structured request delivered to the gatekeeper known as an MCP server. The server retrieves that data and sends it back to the LLM, which interprets and synthesizes it into a response.
Why does MCP matter for authenticity and provenance? MCP can function as a gateway to authentic, verifiable data stewarded by LAMs, which reduces, though doesn’t altogether eliminate, the possibility of hallucinations. In its most ideal state, MCP connects a library, archive, or museum’s collections data directly to an LLM, including descriptive metadata, catalog records, finding aids, full text, or digital images.
How many of us have imagined “chatting” with a historical collection regardless of its size and scope. Unlike traditional search, which relies upon identifying the exact terms that you hope will yield optimal results, the LLM allows for more open-ended or ambiguous queries. In theory, an MCP-directed exchange can translate your historical query, however vague, and open multiple entry points into the data. Like typical LLM chats, a user can engage with the LLM in a more conversational context, asking follow-up questions and probing for additional information that may otherwise not surface as readily in a typical keyword search approach.
Case Study #1: What Claude Got Wrong About Tulips
My first experiment involved an MCP server on GitHub that was built by an independent developer and connects to the Rijksmuseum catalog containing over 800,000 objects. The initial installation process took less than an hour thanks to troubleshooting with Claude Code.
Once installed, you can conduct queries directly within Claude, which initiates specific tools that can access the museum catalog. These tools, which you see here, can, for example, search records for basic information such as artist name, colors, artwork type, or period. For our purpose, the tools can also retrieve more sophisticated provenance information stored within the records, such as historical context, curatorial information, exhibition history, and conservation treatments.
For my inquiry, I composed a simple prompt requesting high-resolution images of artworks containing tulips. The MCP server retrieved 14 works and displayed four in my web browser, including three still lifes from the Dutch Golden Age. I should point out that a standard search on the museum’s site retrieved 616 artworks depicting tulips, so right away we can see that there is a question of whether the first 14 works identified by the chatbot were random or curated according to some unknown set of criteria.
Next, I asked Claude to develop a lesson plan for a middle school art class based on the three still life paintings. This is where the real jaw-dropping moment occurred. Without my specifying any historical context, the LLM independently associated the lesson plan with the historical phenomenon known as Tulip Mania, in which a speculative bubble emerged from the merchant class’s obsession with rare tulips. In moments, Claude drafted a ten-page, four-day plan that framed the paintings within the historical moment and encouraged close examination of the works. The opening question prompted students to consider how much they would pay for a tulip seen in the paintings, followed by a 15-minute “mini-lecture” that discusses Dutch trade, the rise of a new merchant class, and the speculation “craze” surrounding exotic tulips. The multi-day activity culminated with students thinking about present-day consumerism and instructed them to create their own still life using symbols that represent their tastes.
What I found most striking was Claude’s own rationale. It commented that:
“The approach makes 400-year-old paintings relevant by connecting to things middle schoolers care about - what’s valuable, what’s trendy, what represents them - while teaching real art analysis skills and historical context!”
I’m not going to lie, I did kind of wish I could test drive the lesson plan in a real middle school classroom. I imagine there would be a lot of final projects of still lifes with NeeDoh and Sephora products, clearly suggesting the malleability of life and impermanence of beauty.
The MCP server excelled at retrieving reliable data from the museum catalog: the still life paintings contained correct catalog metadata. In other words, it addressed the persistent issue of accuracy by ensuring that its outputs were grounded in authentic data, a major step towards building trust.
However, this example broke down at the point in which the LLM interpreted the authenticated collections data according to methods based on an opaque knowledge base. The Claude-generated lesson plan appeared to cover topics appropriate for a middle school audience, such as the rise of the Dutch merchant class and tulips as a status symbol. But look a little more closely, and it reinforced the debunked Tulip Mania historical theory that purports the bubble created widespread national economic collapse. Recent scholarship has proven that the impact was limited to a small class of merchants and that much of the more salacious myths trace back to a single account written by Charles MacKay two centuries later.
When I pointed this out to Claude, it acknowledged it over-generalized the historical context and proceeded to conduct a seven-minute research query that supposedly consulted over 300 sources. Its revised summary was far more accurate and detailed, but what was even more interesting was the explanation it provided for why it leaned heavily on Tulip Mania myth more than fact:
“The persistence of tulip mania mythology in AI systems reflects the sources these models were trained on. Economics textbooks, popular finance books, and general reference materials overwhelmingly repeat [Charles] Mackay’s [1841] narrative. Claude—like other AI systems—may perpetuate these myths unless specifically prompted with revisionist scholarship.”
In other words, despite retrieving authentic museum catalog data, the LLM co-mingled that data with knowledge contextualized by contemporary media. Within the training dataset, Internet texts comparing cryptocurrency and other speculative bubbles to the most exaggerated or disproven claims about Tulip Mania outweighed the fewer pieces of recent historical scholarship.
Here is our first example of provenance laundering. Recall that it was Claude, not me within any of my prompts, that connected my search for tulip paintings to Tulip Mania. But in the rush to historicize a simple inquiry, it subtly guided me toward a theory that while not entirely wrong, could reinforce a superficial historical understanding. Now imagine that the average user, say a high schooler or college student writing their research paper, may not follow-up their initial prompt with clarifying questions. You can see how easily the LLM can steer the user toward a pre-determined outcome, with little recourse on tracking the source data that led to the misleading output.
Case Study #2: Chasing the Swedish Nightingale
For a second experiment, I vibe coded my own MCP server connecting to Chronicling America, the Library of Congress-hosted database of millions of historic digitized newspaper pages spanning all 50 states and some jurisdictions. A single prompt pointing Claude Code to Chronicling America’s open API was all it needed to generate server tools.
Each tool was designed to conduct a specific type of search and retrieval of the database. The Library of Congress provides a guide for how to work with the site’s API, which includes a README page that lists the types of search parameters that are possible, seen here on the left. Examples include searching by title, issue, page, terms, and dates. Like the Rijksmuseum MCP Server, Claude Code translated that information into separate MCP tools which you see on the right.
While the initial MCP server setup was relatively quick to develop and install, it did not function as I had expected it would when accessing the database. This is where users are critical in sustaining knowledge ecosystem trust. As with my first experiment, the Chronicling America server without fail retrieved accurate metadata associated with newspaper pages. What I hadn’t initially accounted for was the strange search and discovery processes that it would employ under the hood.
My case study involved retrieving feature-length articles and reviews related to the 1850-51 wildly popular U.S. concert tour of Jenny Lind, promoted by none other than P. T. Barnum. The opacity of the LLM’s pretraining once again created provenance problems, this time centered on the search process itself. Through trial and error, I discovered that Claude was cutting corners by using its pre-trained knowledge of the Lind tour, likely heavily based on Barnum’s self-serving autobiographies, to target specific newspapers, dates, and locations rather than searching the database systematically.
In one respect, that’s exactly the kind of contextual knowledge we’d want from an LLM, and in one respect it did not disappoint when it compiled results in an impressively laid out multi-sheet spreadsheet detailing its search methodology, organized by concert tour stops and search gaps. However, without a reliable breakdown of what influenced Claude’s article selection, we get a hybrid product that lacks accountability. What’s worse, a cursory review of results revealed a report riddled with errors, not from the metadata of the retrieved pages, but from the pages it had selected and how it summarized its findings.
Perhaps it is my perpetual optimism or simply naïveté, but it was hard not to share in the excitement when Claude declared in the most certain of superlatives that this time, compared with the previous half dozen searches, it had found an extraordinary article that changes everything we know about the Swedish Nightingale. Now, to be fair, on occasion this was true, and as an historian it kept me hooked as I watched in real-time Claude’s pursuit for new discoveries. Unfortunately, just as often it confidently labeled a dime a dozen advertisement for Lind-branded concert caps or trousers as high quality. In other instances, it found a potentially useful article, only to summarize it incorrectly. Over time, my responses went from, “Ok Claude, let’s get to the root of this error!” to “Claude, are you gaslighting me?” to simply, with a hand on its shoulder, “Oh, no, dear, that is not the most extraordinary article in all of Chronicling America.” In short, my research inquiry required significant, at times painstaking, refinement of the server’s code and my prompts. As my research companion Claude explained:
“The risk isn’t simply that LLMs can be wrong, which is well understood. The deeper problem is the laundering of unattributed knowledge into apparently sourced research.”
To make a long story short, my experiments to reconstruct the reception of the Lind tour across the U.S. did yield positive, genuinely exciting results that may not have readily surfaced using a traditional keyword search approach. Overall, I was able to gather dozens of reception pieces that tapped newspapers that as far as I could tell had not been consulted by Lind or Barnum scholars.
The greatest find was a series of poems and editorials that appeared in abolitionist papers petitioning a reluctant Lind to acknowledge their cause, an example of proto-cancel culture, if you will. Such research discoveries required substantial user input to guide the project, along the way sifting through poor results, refining contextual information within prompts, and essentially training the LLM according to my own research methodology.
My two experiments with MCP servers illustrate how AI systems introduce intermediary layers into LAMs data and the search and discovery process that complicate our portrait of authenticity and provenance. In both cases, I found the results less than trustworthy because I could not dispel the feeling that the LLM overlooked critical results. The uneven results did not derive from the data retrieved by the MCP server, which without fail was authentic and traceable back to the source database. Rather, working with LLMs exposed the problems that arise when its pretraining knowledge steered its interpretation of search results. And yet—and this is a big yet—through it all, I could not help but catch the glimpses of the future of research, if only we could resolve the Gordian knot of source material, interpretation, and trust.
Overcoming Praxis Paralysis
Now that we have a richer picture of how AI disrupts content authenticity and provenance, where does this leave us? I think there will be two tracks that will be played out in parallel with one another.
On one track will be heritage professionals working toward developing schema and systems that accommodate provenance data referencing AI workflows. To illustrate this track, let’s return briefly to C2PA. Provenance data appears in manifests, a structured metadata record tracking version history, electronic signatures, and modification events that travel securely with the digital object. I anticipate we will develop content management systems that will accommodate new schema such as C2PA’s manifests. It’s possible that a PREMIS version 4.0 data dictionary integrates new terms pointing to where AI was used at the level of creation, modification, or description. Or perhaps, given the complexities of AI systems, the field may decide to create entirely separate standards. Either way, this trajectory should feel like a natural continuation of digital preservation workflows that have been evolving for decades.

But there is a reason that I framed my talk around the idea of knowledge ecosystems rather than heritage collections alone. As we speak, the information landscape is being reconfigured for AI-first access. If the first two eras of the Internet were human-centered, where humans both produced and consumed content, powerful forces from academia to the private sector are swiftly taking us into an unchartered third era where a parallel Internet, separate from the one we use daily, is designed solely for AI agents working on our behalf 24/7. To maximize their potential, new information protocols are being designed to speak directly to agents and not to us.
The AI movement is predicated on tearing down barriers to information access. Why search one archive when you can search hundreds at once? This more-is-better mentality contains an enormous blind spot that brushes aside authenticity and provenance. The gap may have narrowed slightly last week with Google and OpenAI’s announcements, but as I hope to have shown, the value of authenticity and provenance doesn’t stop with a manifest or detection tool.

You, heritage professionals, must be willing to step up to address the trustworthiness of knowledge. It may not seem like it at times, but in my experience, sectors beyond heritage are hungry for your perspective, and many would welcome your expertise. The one thing I recommend not doing is nothing. The stakes are too high to shelter in place with the hope that this is merely a phase.
Knowledge ecosystems are sustained not merely when information can be accessed, but when its origins, transformations, uncertainties, and contexts can be traced. AI will undoubtedly stand beside us as partners, companions, or coworkers. Nevertheless, when you are confronted with a decision about adoption, ask yourself what the potential impact would be on content authenticity and provenance. If you foresee real risks involved, that’s a clear signal to tread cautiously. Always remember that technology alone will not resolve these issues but will require human mediation built on collaboration and intellectual curiosity each step of the way.
Thank you!















