Empowering Genealogists with AI: Presentation to National Genealogical Society (NGS)

First-time visitors: If the topic of AI-assisted genealogy is of interest to you, you can subscribe to the right to receive notice of new posts at AI Genealogy Insights. And Welcome! - Steve

I gave a presentation titled “Empowering Genealogists with AI” to the National Genealogical Society on September 6, 2023. The talk covered emerging beneficial use cases of AI-assisted genealogy while addressing current limits, privacy issues, and ethical concerns, and includes a homework assignment you can complete to create an AI-generated handout as practice (jump to 1:04:00 for instructions). You can view the talk here: https://youtu.be/npQaRJbzE1s.

I had a great time with the presentation. I hope you enjoy it. And I am grateful to NGS for the opportunity.

Crafting a Genealogy Prompt for ChatGPT: Five Valuable Components

The post was prepared as part of my upcoming talk: The National Genealogical Society is hosting a MemberConnects! event on Empowering Genealogists with Artificial Intelligence. Join us Wed. 6 Sept 2023, at 8 p.m. ET. Register here: https://bit.ly/NGSMemberConnects6Sept2023

As we venture deeper into the digital age, the intersection of genealogy and artificial intelligence (AI) is becoming increasingly more exciting. Today, I’m diving into a key skill that will empower your genealogical research with AI: how to craft effective prompts for large language models (LLMs) like ChatGPT.

If you’ve been following the ongoing conversation about genealogy and AI, you’ll recall that while LLMs have an arsenal of abilities, they also come with some quirks. To guide them effectively, we need to be thoughtful in our approach.

So, without further ado, let’s discuss the five components that make for a great genealogy prompt!

1. A Role: Begin by imagining you’re recruiting an expert for a specific task. What’s their profession? What expertise should they possess? By defining a role, you’re setting the stage and giving the LLM a context to operate within.

2. A Goal: What do you want to achieve? This is your endgame. Being explicit here ensures that the AI has a clear understanding of your expectations.

3. A Text: Since genealogy is grounded in factual research, provide the text you’d like the AI to process. By supplying the exact text, you mitigate the risk of the AI generating fictitious information (a phenomenon known as “hallucinating”).

4. A Task: Simplify your request. Break it down into manageable steps, just as you’d explain a process to a bright but inexperienced intern. This ensures the LLM knows the sequence of operations it should perform.

5. A Flask: While the term “flask” is playfully chosen for the rhyme, it represents the container or format you’d like your response in. This could range from a narrative report to a structured database table.

For instance, consider this prompt:

PROMPT: You are an expert genealogist and a seasoned data scientist. Your goal is to extract structured relationship data from a provided obituary. Find below an obituary for John Smith from the July 20, 1923, New York Times, page 18. Extract any explicitly stated relationship information; include a quote that supports that determination. Format your response as a CSV file and display that in a code window for easy copy-and-paste.

By following this structure, you provide the LLM with a roadmap to navigate your request. The result? You get precisely the information you’re looking for, formatted just the way you want it.

In conclusion, the digital world of genealogy is at our fingertips, and with tools like ChatGPT, the possibilities are endless. But remember, like any tool, its effectiveness lies in how we use it. With a role, goal, text, task, and flask in hand, you’re all set to harness the full power of LLMs for your genealogical pursuits. Happy researching! 🌳🔍🖥️

A (Possible) Future of A.I. Genealogy Research: Open Archives and ChatGPT

UPDATE: Shortly after sharing this post, Bob Coret, the creator of the “Open Archives” plugin (and a founder of the site), got in touch. The “Open Source” plugin in more impressive than my cursory exploration revealed. The plugin is more thoroughly documented at his blog post, which is linked at bottom along with further observations.

For those engaged in A.I. genealogy and digital explorations, this piece investigates a modest recent progression in A.I.-assisted genealogical research. You may already be familiar with this if you’re a seasoned researcher, but for those who haven’t encountered it yet, this post will introduce a compelling new tool. It’s particularly relevant for those interested in Dutch and Belgian archives, but the implications could extend far beyond these boundaries.

Encountering Open Archives

While exploring the digital landscape for genealogy-related plugins, I made an intriguing discovery: the “Open Archives” plugin. This tool offers a novel approach to accessing genealogical data from Dutch and Belgian archives and societies. Here’s how the plugin describes itself: “Search the genealogical data of Dutch and Belgian archives and societies via Open Archives.”

[Image 1 Caption: A glimpse of the Open Archives plugin interface.]

A preliminary test shows promising results. The plugin appears to effectively search records across these archives, providing links to found records on the original archive websites.

[Image 2 Caption: A sample prompt and response demonstrating the plugin’s capabilities.]

However, it’s worth noting that the plugin currently does not offer additional AI processing of the records within ChatGPT directly. That is, the plugin doesn’t import the record data into the chat session, a feature that could potentially add another layer of convenience and efficiency to our research.

Exploring the ChatGPT Session

To fully understand the potential of this tool, consider this: anyone can access the ChatGPT session via this link, and the returned links are live. Registered ChatGPT users can continue the chat and ChatGPT Plus subscribers can also carry on the conversation with the plugin enabled.

Access the ChatGPT Session

Diving Deeper: Plugin Data Inspection

Beyond simply using the plugin, it’s also possible to examine the data passed to the plugin and returned by it. All you have to do is click the down arrow beside the plugin title. In this case, for example, click the down arrow next to the “Used Open Archives” label.

[Image 3 Caption: The Open Archives plugin’s ‘More Information’ button.]

The Nitty-Gritty: Structured Data

The data exchange between the user and the plugin happens in the form of JSON structured data, a widely used format that allows for easy copying out of the plugin.

[Image 4 Caption: A snapshot of the JSON prompt and response.]

The Road Ahead

While the Open Archives plugin doesn’t yet import the record data into the chat session, this pioneering tool provides a glimpse of what might be possible in the future of digital genealogical research. It paves the way for a more interactive, AI-enhanced exploration of historical records, opening up new avenues for discovery and understanding. As researchers and enthusiasts, let’s keep our eyes on the horizon for what’s coming next!

UPDATE: Shortly after sharing this post, Bob Coret, the creator of the "Open Archives" plugin (and a founder of the site), got in touch. The "Open Source" plugin in more impressive than my cursory exploration revealed. The plugin is more thoroughly documented at his blog post, which is linked below.

* Contrary to a comment in the linked article, the plugin does import record data into the chat session. In simple cases, you can ask the system for birth information (date/place) for a specific person, and then continue in a chat style by asking when the person died.
* A more complex query that can be made is "did they have children". For this, ChatGPT needs to call a specific function using a unique identification (GUID) of a marriage certificate.
* Another impressive query is "how old was {name}". For this, ChatGPT makes two requests to Open Archives (for birth and death records) and calculates the person's age.
* Bob Coret has documented his experiences with the plugin in more detail in a blog post, which is written in Dutch but easily translated to your first language.

https://blogbob.coret.org/2023/06/open-archieven-als-plugin-voor-chatgpt.html

AI Genealogy Use Case Guide: How-to Get from Story to Structured Data, 2: Create GEDCOM (family tree) files from obits, articles, & announcements

Introduction:

  • In the world of genealogy research, information is scattered across various sources, including narrative texts such as birth and wedding announcements, obituaries, and newspaper articles. These unstructured narratives can be challenging to manage and analyze. In this blog post, we will explore the specific task of quickly extracting valuable data such as names, relationships, dates, and places from these texts and converting them into structured data formats such as GEDCOM files, the format used to build and exchange family trees and to share genealogical data. This data extraction and storage process is tremendously beneficial for for streamlining genealogical research and making connections between family members, ancestors, and historical events.
  • To accomplish this task, we will be utilizing ChatGPT, an advanced AI language model developed by OpenAI. ChatGPT is capable of processing and extracting information from large volumes of text, making it an ideal tool for genealogy enthusiasts seeking to organize and analyze their research data efficiently.

Objectives:

  • Automate the extraction process: Utilize ChatGPT to efficiently extract names, relationships, dates, and places from various sources, such as announcements, obituaries, and newspaper articles, minimizing manual effort and speeding up the research process.
  • Improve data organization: Convert the extracted information into structured data formats such as GEDCOM files to facilitate better organization, storage, and retrieval of genealogical data.
  • Enhance data analysis: Enable genealogy researchers to analyze structured data more effectively, identify patterns, and uncover hidden connections between family members and historical events.
  • Save time and resources: Streamline the research process by reducing the time spent on manually extracting and organizing data, freeing up more time for analysis and interpretation.
  • Increase research accuracy: Minimize human errors and inconsistencies in data extraction and organization by leveraging ChatGPT’s advanced language processing capabilities.

Requirements:

  • Access to AI: Obtain a free or paid subscription to an artificial intelligence service such as OpenAI’s ChatGPT. Other AI options include Google’s Bard, Anthropic’s Claude, or Microsoft’s Bing Chat, but in April 2023, OpenAI’s GPT-4 based ChatGPT is strongest.
  • Input data: Provide ChatGPT with text from sources such as birth or wedding announcements, obituaries, and newspaper articles, containing information about names, relationships, dates, and places relevant to genealogical research.
  • Family tree software to read and use the GEDCOM file created here. Most genealogy applications and website and utilize GEDCOM files to share family trees and genealogical data; these include desktop applications such as RootsMagic, Family Tree Maker, Gramps, and online resources such as Ancestry, DNA Painter, MyHeritage, and FindMyPast.
  • Optional: Formatting requirements: Ensure that the input text is free of major errors or inconsistencies. Although ChatGPT can handle some level of noise in the data, better-formatted input will yield more accurate and reliable results. Our earlier Use Case Guide on Cleaning OCR Text quickly steps you through this process.
  • Helpful: Genealogy research resources: Familiarize yourself with various genealogical research methods, repositories, and databases to: (1) know where to find the texts to data mine, and (2) effectively contextualize and validate the information extracted by ChatGPT.

Caveats for the Careful on Large Language Models in Genealogy (April 2023):

  • No live internet access, with data current only up to September 2021
  • Unreliable for fact-based research, relying on statistical language patterns
  • Official ChatGPT warning: may produce inaccurate information
  • Mainly used for information processing, not discovering new data
  • Limited to 1500 words input/output (approx. 4k tokens)
  • Chatbots lack traditional memory, necessitating careful management of conversation
  • See Use Case Guide #1: Cleaning OCR Text for detailed information on each caveat above.

How To: Methodology:

  • Step 1: Get a free or paid AI account. In April 2023, OpenAI’s GPT-4 based ChatGPT is strongest, but the free version based on GPT-3.5 will also work; you can get a free account at https://chat.openai.com/auth/login.
  • Step 2: Find and prepare your input text. In spring 2023, most publicly-accessibly AI systems are based on large language models that are untethered to reality or knowledge systems; they work by selecting the next statistically most likely word based on your prompt and previous utterances in your chat. For this reason, fact-based researchers such as genealogists restrict the AI to working only on the data you input. In this series of Use Case Guides, we have been using texts from publicly available sources such as the Chronicling America newspaper archive from The Library of Congress and the National Endowment for the Humanities. Our Use Case Guide #1: How to Clean Raw and Poor OCR Text breaks-down this step down in a detailed walk-through; refer to that guide if you would like help with this step. This Guide uses an obituary first published 100 years ago this week in one of my state’s capital newspapers: “Westmoreland Club Honors J. E. Royall,” Richmond Planet (Richmond, VA) 1883-1938, April 21, 1923, Page 8, Image 8; Image and text provided by Library of Virginia; Richmond, VA; < https://chroniclingamerica.loc.gov/lccn/sn84025841/1923-04-21/ed-1/seq-8/ > [accessed: 17 April 2023]. I processed and cleaned the raw (nightmarish) Chronicling American text using the steps in Usage Guide #1.
  • Step 3: Start a new ChatGPT session. It is important to start a new ChatGPT session when beginning a new genealogical task because (lacking both short-term and long-term memory) the chatbot re-ingests up to the previous 400 lines of your “dialogue” in order to simulate a conversation; this can have the unintended effect of contaminating your chat with information from pervious utterances in the current session. See “Don’t Get Burned by Spicy Autocomplete” for more information about this concern.
  • Step 4: Write your prompt. Your prompt will include two parts: (a) your instructions to the AI, and (b) the input text from which you want to extract structured genealogical data. We’ll discuss both parts in turn.
  • Step 4(a): Write your instructions to the AI. The instructional component of a prompt itself has subcomponents. Here you can see that the first part of the instructions are directing the AI to assume a role, in this case, of a genealogist; this has the effect of providing a context for the AI’s response. Next, action verbs to direct the AI; in this case “Find,” “Prioritize,” “Create,” “Include,” and “Respond.” These will change depending on your task; if you have troubling crafting this part, ChatGPT can help: write the instructions as best you can, then use ChatGPT to “convert these statements to the imperative mood“; this has the effect of changing your statements to the desired form “[you, the AI] do (verb) this.” Finally, you will see that here we are directing the AI to create a table of data; this is the most simple form of structured data, and perhaps the most meaningful and accessible to the human genealogist! Later, I’ll show you how to transform this table data into forms more suited for genealogical tools such as tree making applications, spreadsheets, and databases. To assist in the verification of our work, the AI is instructed to show its work, that is, to include the evidence it used to make a relationship determination by quoting the passage it relied upon to state a relationship. So far, thus prompted, contained, restrained, and instructed, I have not witnessed a fabrication or hallucination of a relationship. If you do, capture and save your whole session; I’d love to see it.
PROMPT: Assume the role of an expert, professional genealogist. Find below the text of an obituary. Prioritize fidelity to the information below. Create a table of named relatives of the deceased. Include only explicitly named relationships. Respond in the form of a markdown table with the column headings: Deceased | Person 2 | Relationship | Evidence (where evidence is the quoted text from the article used to determine relationship).
  • Step 4(b): Paste your input text below the instructions. Below your instruction, paste the text from which you would like to AI to process. Remember, as of April 2023, we are limited to about 1500 words input. Folks may tell you that you can upload more, perhaps by asking the AI to accept your input in parts, and it will agree to do that, but if you upload more than about 400 lines or about 1500 words, the AI will drop and ignore parts of your input. (In OpenAI’s technical jargon, we are limited to 4096 “tokens,” more akin to syllables than words, but for simplicities sake, about 1500 words input and 1500 words output.) [NOTE: If you need to enter a newline or line break in the ChatGPT edit box, Shift-Enter will give you a new line without submitting the request.]
  • Step 5: Examine your results; adjust as needed. Just as writing means re-writing, so prompt engineering means prompting and re-prompting. I never get the best results on a first attempt, so I expect to refine a prompt until the AI is producing the data in the form I want.

The results here are exactly as expected. This use case is a greater accomplishment than may be apparent to some. The imagination or understanding of how this use case will soon scale (become more powerful) is sometimes the missing piece. Here, eight relationships were extracted from a 700-word obituary; that is admittedly weak tea. But in time we will be able to process book chapters (50-pages announced already), whole books (this year or 2024), and entire archives after that. That’s the big deal that’s coming.

  • Step 6: Set the context for your GEDCOM request prompt. Set the context of the GEDCOM request by asking the AI about its familiarity with GEDCOM files.
PROMPT: Are you familiar with the GEDCOM file format and standard?
  • Step 7: Prompt for the creation of the GEDCOM file. There are several items to note about this prompt. First, we are directing to AI to transform the table of named relationship created earlier; that table included direct quotations from the source article for ease of validation and verification, but since that is not desired in the GEDCOM file, we instruct the artificial intelligence to omit that column of information. We do want the source of this information included with the GEDCOM file, so we supply that here to ChatGPT.
PROMPT: Create a GEDCOM file from the table of named relatives above. Omit "Evidence" column. Include source information: "Westmoreland Club Honors J. E. Royall," Richmond Planet, Richmond, Va. 1883-1938, April 21, 1923, Image 8, Image and text provided by Library of Virginia; Richmond, VA
Persistent link: https://chroniclingamerica.loc.gov/lccn/sn84025841/1923-04-21/ed-1/seq-8/

GEDCOM FILE:
  • Step 8: Examine your results. I have found that ChatGPT very reliably creates accurate and functional GEDCOM files using this complete method. Other artificial intelligences may not be as reliable: Anthropic’s Claude will create an accurate GEDCOM file, but fail to add the newlines (carriage returns or hard line breaks) needed to create a functional GEDCOM; asking again for those newline characters usually works; I have not been able to successfully create an accurate and functional GEDCOM with Google’s Bard, Perplexity AI, nor Microsoft’s Bing Chat.
  • Step 9: Save your results. You need to save your GEDCOM file as a text file. This means finding and using your computer’s text file editor. The basic Windows text editor is Notepad, so Windows users will open Notepad with a new, blank file. Then, in the ChatGPT code window, click the “Copy code” link at the upper right corner of the code window. Switch to the blank text file, paste the GEDCOM data into the text file, and save the file with a name such as “Royall.ged”; the “.ged” extension (last characters of the file name) is important. Remembering to include this file name extension will enable your genealogical apps and sites to recognize this text file as a family tree file.
  • Step 10: Open, test, verify, and confirm your work. At this point, you can open your genealogy application such as RootsMagic, Gramps, or Family Tree Maker and open or import the GEDCOM file (you will need to check that application’s instructions for opening and/or importing a GEDCOM file). You will usually want to open the GEDCOM file as a new tree, as opposed to merging it into an existing tree. Compare the information now in your new family tree to the information stated in the birth or wedding announcement, obituary, or newspaper article.

Results and Analysis:

  • Expected Outcomes: By using AI for this genealogy task, you can expect the extraction of key information such as names, relationships, dates, and places from various text sources, subsequently saving the data in a structured format such as GEDCOM files. This will facilitate easier sharing of family trees and the exchange of genealogical information.
  • Accuracy and Reliability: While ChatGPT is a powerful AI model, the accuracy and reliability of the results will depend on the quality of the input data and the clarity of the information present. In most cases, ChatGPT can accurately extract relevant data points, but manual review and validation is required to ensure the information is consistent with your research goals.

Conclusion:

  • In conclusion, using ChatGPT for extracting structured data from narrative sources like birth and wedding announcements, obituaries, and newspaper articles offers significant benefits and some limitations. The technology has the potential to greatly enhance genealogy research by automating the extraction of names, relationships, dates, and places, saving time and effort for researchers. The ability to convert this information into a structured data format such GEDCOM files further streamlines the research process and facilitates data organization, the sharing of family trees and the exchange of genealogical information.
  • However, limitations need to be considered. ChatGPT’s accuracy may vary depending on the quality of the input text, especially if dealing with raw OCR text or handwritten documents. Additionally, ChatGPT might struggle with complex relationships and ambiguous information present in the narratives. To overcome these challenges, AI Genealogists need to manually review and verify the extracted data.
  • For further improvement and exploration of AI in genealogy, researchers should consider integrating ChatGPT with other natural language processing tools or specialized genealogy software to enhance its capabilities. Collaborating with AI developers to create tailored models for genealogy research could further optimize the extraction process and improve overall accuracy. Encouraging users to share their experiences and provide feedback will contribute to the ongoing refinement of AI solutions for genealogy.
  • Ultimately, employing AI tools like ChatGPT for genealogy research has the potential to revolutionize the field, making it more accessible, efficient, and accurate. As AI technology continues to evolve, the possibilities for its application in genealogy will only expand, benefiting researchers and family historians alike.

Call to Action:

  • Try ChatGPT for your genealogy tasks: We encourage you to harness the power of ChatGPT for your genealogy research. Experience firsthand the benefits of using AI to extract structured data from narrative sources like birth and wedding announcements, obituaries, and newspaper articles.
  • Share your experiences: We would love to hear about your experiences using ChatGPT for genealogy tasks. Share your successes, challenges, or any interesting insights you’ve gained through utilizing AI in your research. Your feedback can help improve the technology and benefit the entire genealogy community.
  • Ask questions and seek advice: If you have any questions or need assistance with using ChatGPT for genealogy tasks, feel free to post them in the comments section below or reach out to us on social media. Our community of experts and fellow genealogy enthusiasts will be more than happy to help.
  • Connect with others and expand your knowledge: Join genealogy forums, social media groups, and other online communities where you can connect with others who are using AI for genealogy research. These platforms are excellent resources for sharing tips, tricks, and best practices, as well as staying up-to-date with the latest advancements in AI technology.
  • Explore additional resources: To further enhance your understanding of AI in genealogy and to make the most out of ChatGPT, check out the provided links to tutorials, support forums, and related articles. Continuously learning and staying informed will help you maximize the potential of AI in your genealogy research.

AI Genealogy Use Case Guide: How-to Get from Story to Structured Data, 1: from Text to Table Data, from Stories to CSV files

Introduction:

  • In the world of genealogy research, information is scattered across various sources, including narrative texts such as birth and wedding announcements, obituaries, and newspaper articles. These unstructured narratives can be challenging to manage and analyze. In this blog post, we will explore the specific task of quickly extracting valuable data such as names, relationships, dates, and places from these texts and converting them into structured data formats such as tables, JSON files, and near-universally usable CSV files (great for importing into spreadsheets such as Excel and Google Sheets and into databases such as MySQL and AirTable). This data extraction and storage process is tremendously beneficial for for streamlining genealogical research and making connections between family members, ancestors, and historical events.
  • To accomplish this task, we will be utilizing ChatGPT, an advanced AI language model developed by OpenAI. ChatGPT is capable of processing and extracting information from large volumes of text, making it an ideal tool for genealogy enthusiasts seeking to organize and analyze their research data efficiently. Stay tuned as we dive into the objectives, requirements, and methodology of using ChatGPT for genealogy data extraction and organization.

Objectives:

  1. Automate the extraction process: Utilize ChatGPT to efficiently extract names, relationships, dates, and places from various sources, such as announcements, obituaries, and newspaper articles, minimizing manual effort and speeding up the research process.
  2. Improve data organization: Convert the extracted information into structured data formats (e.g., tables, JSON, and CSV files) to facilitate better organization, storage, and retrieval of genealogical data.
  3. Enhance data analysis: Enable genealogy researchers to analyze structured data more effectively, identify patterns, and uncover hidden connections between family members and historical events.
  4. Save time and resources: Streamline the research process by reducing the time spent on manually extracting and organizing data, freeing up more time for analysis and interpretation.
  5. Increase research accuracy: Minimize human errors and inconsistencies in data extraction and organization by leveraging ChatGPT’s advanced language processing capabilities.

Requirements:

  1. Access to AI: Obtain a free or paid subscription to an artificial intelligence service such as OpenAI’s ChatGPT. Other AI options include Google’s Bard, Anthropic’s Claude, or Microsoft’s Bing Chat, but in April 2023, OpenAI’s GPT-4 based ChatGPT is strongest.
  2. Input data: Provide ChatGPT with text from sources such as birth or wedding announcements, obituaries, and newspaper articles, containing information about names, relationships, dates, and places relevant to genealogical research.
  3. Optional: Formatting requirements: Ensure that the input text is free of major errors or inconsistencies. Although ChatGPT can handle some level of noise in the data, better-formatted input will yield more accurate and reliable results. Our earlier Use Case Guide on Cleaning OCR Text quickly steps you through this process.
  4. Optional: Data storage and processing tools: Utilize software and tools like Microsoft Excel, a CSV editor, or a MySQL client to store, manage, and analyze the structured data extracted by ChatGPT.
  5. Helpful: Genealogy research resources: Familiarize yourself with various genealogical research methods, repositories, and databases to: (1) know where to find the texts to data mine, and (2) effectively contextualize and validate the information extracted by ChatGPT.

Caveats for the Careful on Large Language Models in Genealogy (April 2023):

  • No live internet access, with data current only up to September 2021
  • Unreliable for fact-based research, relying on statistical language patterns
  • Official ChatGPT warning: may produce inaccurate information
  • Mainly used for information processing, not discovering new data
  • Limited to 1500 words input/output (approx. 4k tokens)
  • Chatbots lack traditional memory, necessitating careful management of conversation
  • See Use Case Guide #1: Cleaning OCR Text for detailed information on each caveat above.

Caveats for the Bold

  • These initial Use Cases are admittedly weak tea: limited and narrow in function and capacity; these constraints reflect the abilities and token limits of AI systems for fact-based research in April 2023.
  • For now, think “Lego pieces” not “Post-Doc Assistant”; that is, in spring 2023, don’t imagine AI is a magic genie that can do all your work for you, like a post-doc assistant; instead, AI-assisted genealogical tasks are now more like a growing Swiss army knife or set of Lego blocks with which you can build tools to solve larger problems. It doesn’t take too much creativity to imagine how even these modest use cases can today be linked/chained and combined with each other to accomplish larger genealogical goals; soon enough, I imagine, the larger goals will be one-step AI-assisted tasks. But, for now, enjoy playing with the fundamental building blocks of more powerful systems to come.

How To: Methodology:

  • Step 1: Get a free or paid AI account. In April 2023, OpenAI’s GPT-4 based ChatGPT is strongest, but the free version based on GPT-3.5 will also work; you can get a free account at https://chat.openai.com/auth/login.
  • Step 2: Find and prepare your input text. In spring 2023, most publicly-accessibly AI systems are based on large language models that are untethered to reality or knowledge systems; they work by selecting the next statistically most likely word based on your prompt and previous utterances in your chat. For this reason, fact-based researchers such as genealogists restrict the AI to working only on the data you input. In this series of Use Case Guides, we have been using texts from publicly available sources such as the Chronicling America newspaper archive from The Library of Congress and the National Endowment for the Humanities. Our Use Case Guide #1: How to Clean Raw and Poor OCR Text breaks-down this step down in a detailed walk-through; refer to that guide if you would like help with this step. This Guide uses an obituary first published 100 years ago this week in one of my state’s capital newspapers: “Westmoreland Club Honors J. E. Royall,” Richmond Planet (Richmond, VA) 1883-1938, April 21, 1923, Page 8, Image 8; Image and text provided by Library of Virginia; Richmond, VA; < https://chroniclingamerica.loc.gov/lccn/sn84025841/1923-04-21/ed-1/seq-8/ > [accessed: 17 April 2023]. I processed and cleaned the raw (nightmarish) Chronicling American text using the steps in Usage Guide #1.
  • Step 3: Start a new ChatGPT session. It is important to start a new ChatGPT session when beginning a new genealogical task because (lacking both short-term and long-term memory) the chatbot re-ingests up to the previous 400 lines of your “dialogue” in order to simulate a conversation; this can have the unintended effect of contaminating your chat with information from pervious utterances in the current session. See “Don’t Get Burned by Spicy Autocomplete” for more information about this concern.
  • Step 4: Write your prompt. Your prompt will include two parts: (a) your instructions to the AI, and (b) the input text from which you want to extract structured genealogical data. We’ll discuss both parts in turn.
  • Step 4(a): Write your instructions to the AI. The instructional component of a prompt itself has subcomponents. Here you can see that the first part of the instructions are directing the AI to assume a role, in this case, of a genealogist; this has the effect of providing a context for the AI’s response. Next, action verbs to direct the AI; in this case “Find,” “Prioritize,” “Create,” “Include,” and “Respond.” These will change depending on your task; if you have troubling crafting this part, ChatGPT can help: write the instructions as best you can, then use ChatGPT to “convert these statements to the imperative mood“; this has the effect of changing your statements to the desired form “[you, the AI] do (verb) this.” Finally, you will see that here we are directing the AI to create a table of data; this is the most simple form of structured data, and perhaps the most meaningful and accessible to the human genealogist! Later, I’ll show you how to transform this table data into forms more suited for genealogical tools such as tree making applications, spreadsheets, and databases. To assist in the verification of our work, the AI is instructed to show its work, that is, to include the evidence it used to make a relationship determination by quoting the passage it relied upon to state a relationship. So far, thus prompted, contained, restrained, and instructed, I have not witnessed a fabrication or hallucination of a relationship. If you do, capture and save your whole session; I’d love to see it.
PROMPT: Assume the role of an expert, professional genealogist. Find below the text of an obituary. Prioritize fidelity to the information below. Create a table of named relatives of the deceased. Include only explicitly named relationships. Respond in the form of a markdown table with the column headings: Deceased | Person 2 | Relationship | Evidence (where evidence is the quoted text from the article used to determine relationship).
  • Step 4(b): Paste your input text below the instructions. Below your instruction, paste the text from which you would like to AI to process. Remember, as of April 2023, we are limited to about 1500 words input. Folks may tell you that you can upload more, perhaps by asking the AI to accept your input in parts, and it will agree to do that, but if you upload more than about 400 lines or about 1500 words, the AI will drop and ignore parts of your input. (In OpenAI’s technical jargon, we are limited to 4096 “tokens,” more akin to syllables than words, but for simplicities sake, about 1500 words input and 1500 words output.) [NOTE: If you need to enter a newline or line break in the ChatGPT edit box, Shift-Enter will give you a new line without submitting the request.]
  • Step 5: Examine your results; adjust as needed. Just as writing means re-writing, so prompt engineering means prompting and re-prompting. I never get the best results on a first attempt, so I expect to refine a prompt until the AI is producing the data in the form I want.

The results here are exactly as expected. This use case is a greater accomplishment than may be apparent to some. The imagination or understanding of how this use case will soon scale (become more powerful) is sometimes the missing piece. Here, eight relationships were extracted from a 700-word obituary; that is admittedly weak tea. But in time we will be able to process book chapters (50-page capacity announced already by OpenAI), whole books (this year or 2024), and entire archives after that. That’s the big deal that’s coming.

  • Step 6: Save your work. Save both your ChatGPT session and save your response to a text file. You can now easily download your entire ChatGPT history. You may also want to copy-and-paste the table data to a local file; some of the formatting will be lost if you paste into a plain text file, but pasting into a Word or Google Docs file will preserve the markdown formatting, if that is important to you.
  • Step 7: Wring further data from the text. Named relationships are not the only data that ChatGPT can extract from a text. ChatGPT excels at FAN processing of a text (finding friends, associates, and neighbors that are mentioned in a text). People (“entities” in AI jargon) are not the only data that can be extracted. ChatGPT will also extract places, events, and dates from a text. For example, after extracting the explicit relationships from the obituary, I instructed ChatGPT to extract all named associates from the text:
PROMPT: Create a table of named associates of the deceased; broaden the meaning of associates as wide as possible to include ALL named people in the obituary if their relationship or function at funeral is stated. Respond in the form of a markdown table with the column headings: Deceased | Person 2 | Relationship.
  • Step 8: Create derivative data structures and formats. You can now instruct ChatGPT to create alternate file types such as CSV (common separated values) files which are nearly universally usable by spreadsheets (Excel, Google Sheets), databases (MySQL, MS Access, AirTable), and word processors (Word, Google Sheets). For the technically inclined, GPT-4 is able to convert the table data to JSON files for processing by web applications and custom programming scripts such as Python. In the next Usage Guide, I’ll show you step-by-step how to create a GEDCOM file, used widely to create family trees and exchange genealogical data. Here you can see all how the named associates of the deceased may quickly be downloaded as a CSV file:
  • Step 9: As a last step, ask the AI what you forgot. This is always fun, and reveals that while I may be focused on one type or piece of information, the AI may help me uncover the missing piece to solve a brick wall that was under my nose but which I’d overlooked.
PROMPT: What other genealogically relevant information might I also extract as structed data from this obituary?

Results and Analysis:

  • Expected Outcomes: By using AI for this genealogy task, you can expect the extraction of key information such as names, relationships, dates, and places from various text sources, subsequently saving the data in structured formats like table data, JSON, and CSV files. This will facilitate easier data analysis and integration into your genealogical research.
  • Accuracy and Reliability: While ChatGPT is a powerful AI model, the accuracy and reliability of the results will depend on the quality of the input data and the clarity of the information present. In most cases, ChatGPT can accurately extract relevant data points, but manual review and validation is required to ensure the information is consistent with your research goals.

Conclusions:

  • In conclusion, using ChatGPT for extracting structured data from narrative sources like birth and wedding announcements, obituaries, and newspaper articles offers significant benefits and some limitations. The technology has the potential to greatly enhance genealogy research by automating the extraction of names, relationships, dates, and places, saving time and effort for researchers. The ability to convert this information into structured data formats such as table data, JSON, and CSV files further streamlines the research process and facilitates data organization.
  • However, limitations need to be considered. ChatGPT’s accuracy may vary depending on the quality of the input text, especially if dealing with raw OCR text or handwritten documents. Additionally, ChatGPT might struggle with complex relationships and ambiguous information present in the narratives. To overcome these challenges, AI Genealogists need to manually review and verify the extracted data.
  • For further improvement and exploration of AI in genealogy, researchers should consider integrating ChatGPT with other natural language processing tools or specialized genealogy software to enhance its capabilities. Collaborating with AI developers to create tailored models for genealogy research could further optimize the extraction process and improve overall accuracy. Encouraging users to share their experiences and provide feedback will contribute to the ongoing refinement of AI solutions for genealogy.
  • Ultimately, employing AI tools like ChatGPT for genealogy research has the potential to revolutionize the field, making it more accessible, efficient, and accurate. As AI technology continues to evolve, the possibilities for its application in genealogy will only expand, benefiting researchers and family historians alike.

Calls to Action:

  • Try ChatGPT for your genealogy tasks: We encourage you to harness the power of ChatGPT for your genealogy research. Experience firsthand the benefits of using AI to extract structured data from narrative sources like birth and wedding announcements, obituaries, and newspaper articles.
  • Share your experiences: We would love to hear about your experiences using ChatGPT for genealogy tasks. Share your successes, challenges, or any interesting insights you’ve gained through utilizing AI in your research. Your feedback can help improve the technology and benefit the entire genealogy community.
  • Ask questions and seek advice: If you have any questions or need assistance with using ChatGPT for genealogy tasks, feel free to post them in the comments section below or reach out to us on social media. Our community of experts and fellow genealogy enthusiasts will be more than happy to help.
  • Connect with others and expand your knowledge: Join genealogy forums, social media groups, and other online communities where you can connect with others who are using AI for genealogy research. These platforms are excellent resources for sharing tips, tricks, and best practices, as well as staying up-to-date with the latest advancements in AI technology.
  • Explore additional resources: To further enhance your understanding of AI in genealogy and to make the most out of ChatGPT, check out the provided links to tutorials, support forums, and related articles. Continuously learning and staying informed will help you maximize the potential of AI in your genealogy research.

Genealogy and Artificial Intelligence: Falling Off the Dunning-Kruger Cliff

What do we call the disappointment that first-time users feel when AI tools fail their expectations?

I know I got burned last year and I had my own WTF moment. I certainly don’t blame folks for feeling misled. For SO many reasons. The most harmful may be the hype that likely draws many first-time users–that AI will be a “Google-killer.” Well, since “Google” is synonymous with search the way Xerox was with copiers and Kleenex is with tissues, I don’t blame folks for expecting that the Google-killer would be good at search and research. But as we all soon discover, that’s not what this tool is–at least not in spring 2023.

The AI genealogy learning curve is harsh, and for most of us, involves an early fall off the cliff of the Dunning–Kruger effect. I hope folks recover from that first disappointment, however, to learn how AI can be used in AI genealogy now, in smaller, constrained, less grandiose ways, than we perhaps all expected during our first ChatGPT sessions.

And today’s failed expectations may likely be tomorrow’s breakthroughs. Which is exciting for some, but not unreasonably stressful for others. Hype aside, the pace in which new AI tools and abilities are coming will be quick. So it’s not unreasonable for some to think, “Why should I spend 10 or 100 hours now learning a new skill set that may be obsolete in three years or even three months; I think I’ll wait till the dust settles, and learn the easy version then.”

The concern of replacement among professional knowledge workers, also, is not illegitimate. In a field perhaps tangential to genealogy–law–I heard an interesting discussion among folks more informed than myself. And their discussion went something like this: “AI will not replace lawyers. But AI-savvy lawyers will replace lawyers who aren’t.” Then they imagined a large law firm that currently hires a number of young attorneys each year out of law school. If a fewer number of AI-savvy lawyers can do the work of other lawyers, that could create disruption. Folks near the end of their careers and those with greatest power might not be too concerned. But many others will be.

That said, I’m excited and optimistic about AI genealogy. I hope others will shake off their first disappointments–we’ve all been there–and will ignore the hype and hustle. Learn what AI genealogy can and cannot do today, keeping in mind that those things will change, some more quickly than others. And embrace the thrill of being a life-long learner.

Best wishes, Steve

Hello world!

Welcome to AI Genealogy Insights, where we explore how artificial intelligence can assist genealogists and family history researchers, with a particular focus on:

  • discovering the advantages and limitations of AI,
  • and how genealogists can apply this knowledge.

As someone with training and a background in applied linguistics (natural language processing and computation linguistics–foundations of artificial intelligence), language and literature, computers and programming, writing and storytelling, and genealogy and family history, I am passionate about combining these interests to enhance genealogical research and discovery.

My focus has locked onto the fascinating world of artificial intelligence and its applications in the field of genealogy. Here, you’ll find discussions on what AI is, what genealogists need to know about the current state of AI-assisted genealogy, what currently are good applications of AI genealogy and what are not, how to leverage AI to be more productive, what AI genealogy developments may be coming soon, where useful AI tools and services can be found, and explorations as the limits and boundaries of what is possible with AI genealogy are expanded.

I invite you to explore the site and engage with the ideas presented, as we embark on this exciting journey of harnessing the power of AI to uncover the rich tapestry of our family histories and how we can do more with AI genealogy. There are two areas of the site in which you might begin exploring: Use Case Guides and/or Blog Posts.

The Use Case Guides walk users step-by-step through currently working and proven AI-assisted genealogical tasks. AI Genealogy is a fact-based research and reporting discipline. Yet the large language models powering today’s artificial intelligence (AI) systems weak with facts though great with words. Their weakness with facts and research, however, hardly renders them useless to genealogists. Knowing the benefits and limits of artificial intelligence empowers genealogists today and help us recognize breakthroughs tomorrow. The use cases here illustrate how researchers are successfully using AI Genealogy today.

Another place to begin is by reading a Blog Post that is of interest to you. Blog post topics are more wide-ranging, and may include: AI Genealogy industry news and commentary; recommendations for books, tutorials, podcasts, video, and the like; AI Genealogy Tips and Tricks; and both upcoming and speculative AI Genealogy Use Cases. Blog posts written before April 7, 2023, were originally published on my family history site, Ashe Ancestors. I will continue to post about family history there, but my AI Genealogy research and writing will be shared here.

AI Genealogy: Text to GEDCOMs: Surprises, Cautions, Discoveries

I had an opportunity today to experiment a bit more with using artificial intelligence to create family trees (GEDCOM files) from narrative texts. My goal was to see how much I could limit the AI’s creativity to insert information into the GEDCOM file that wasn’t in my prompt. (Earlier, I had discovered two constraints that are helpful: (1) starting a new chat session before creating GEDCOM files, and (2) including in the prompt the instruction to use only facts given to it to generate the GEDCOM. I discovered today that there are other levers and dials that can more or less control the creative aspects of the AI when creating family trees/GEDCOM files. It was an interesting morning.

First, a note about the AI tool I was using today. Many folks are familiar with the web interface of OpenAI’s ChatGPT; that wasn’t what I was using today. Instead, I was using OpenAI’s “Playground,” an different web interface that you can access when you sign-up for API access, used for web programming. The OpenAI Playground allows you to change several settings (“parameters” in their jargon), which effects what the AI generates. While working the in chat mode with which we’re familiar, I was experimenting today with these settings:

  • Model: akin to personalities with different capabilities; for example, DALL-E will generate images, Whisper will convert audio to text, and GPT-3, GPT-3.5, and GPT-4 that can generate natural language responses (or computer code) given instructions in natural language (these last models are the chatbots with which we’ve become familiar).
  • Temperature: akin to creativity, set on a scale from 0.0 to 1.0.
  • Maximum length: controls the total amount of input and output together (measured in “tokens,” akin to syllables or words); for example, roughly, if the “Maximum length” is set to 1000 words, and you input 750 words, then your output will stop at 250 words; used to control both verbosity and expense (e.g., gpt-3.5-turbo costs $0.002 / 1K tokens); my total cost today were $0.20, or 20 cents, for several hours of experimentation.

I didn’t experiment too much with the “Maximum length” today, as I was mostly interested in learning about how the Model and Temperature settings effected the outcome of a prompt. So I started with this prompt to create a family tree (GEDCOM file) for a family with which you may be familiar, though, as we discovered earlier, you could also use the OCR’d text of an obituary, wedding announcement, or newspaper article. Today, however, I wanted to use something very simple, to focus on some narrow adjustments:

SYSTEM: You are an expert, professional genealogist and computer programmer. Respond only in the form of a GEDCOM file.

USER: Adam and Eve are the father and mother of Cain and Abel.

The SYSTEM instruction gives the AI a bit of instruction about how to respond. For example, if you suggest “SYSTEM: Respond in the voice of William Shakespeare,” then the AI will respond to you in Elizabethan English in iambic pentameter. I instructed the AI to respond only with the text and format used to create family trees, GEDCOM files. And I was pleasantly surprised that that was the way the AI responded. There were some unexpected results, however, in the response:

(click to enlarge)

Can you spot the unexpected (and undesired) hallucinatory information inserted into the GEDCOM? Here it is a bit closer; these are the lines for Eve and Cain:

0 @I2@ INDI
1 NAME Eve / /
1 SEX F
1 BIRT
2 DATE ABT -4000
2 PLAC Garden of Eden
1 FAMS @F1@

0 @I3@ INDI
1 NAME Cain / /
1 SEX M
1 BIRT
2 DATE ABT -3970
2 PLAC Garden of Eden
1 FAMC @F1@

There are two or three pieces of information in the GEDCOM that were not in my input: dates and places. The AI helpfully (or not!) suggested a date and place for Eve’s birth (creation?), about 6000 years ago, in the Garden of Eden, where the AI also places Cain’s birth at 3970 BCE. (Or BC? I didn’t ask, and I’m not sure what the GEDCOM standard states.) And there’s another unexpected hallucination I’ll point out a bit later.

These results were generated with the default settings of the AI. But we need not despair or abandon the project. There are about a half-dozen adjustments we can make to constrain and guide the AI to generate a desired result, using only the information we provide to create a more accurate family tree (GEDCOM file) from some text.

Changing Settings to Reduce Hallucinations

Look more closely at the Model and Temperature at the default settings:

First, I was not using the most recent OpenAI Model, GPT-4.
Second, the “creativity” was cranked up to 0.7, or 70% of its maximum.
Let’s (1) change the model to the best currently available, GPT-4, and (2) turn the Temperature down to 0.0 to try to squelch the creativity and hallucinations. Making those two changes (keeping the prompt and everything else the same), here are the results:

0 @I2@ INDI
1 NAME Eve
1 SEX F
1 FAMS @F1@

0 @I3@ INDI
1 NAME Cain
1 SEX M
1 FAMC @F1@

Excellent. Making those changes (1) to the most recent model and (2) lowering the creativity was successful in eliminating any reference to date or place of birth, which I hadn’t included in my initial prompt. This is good news.

Great. But Not Perfect. Yet.

But, do you see that there is still some information in the GEDCOM that perhaps shouldn’t be there? I perhaps identified the sex or gender of Adam and Eve by describing them as a father and mother, but I never stated in my prompt the gender of their children. Nevertheless, GPT-4 reported the sex of Cain as male, which I hadn’t provided. So I asked it why. Its response made me smile, for a couple of reasons:

(click to enlarge)
USER:
Not that you are wrong, but why did you assign 'male' to Cain and Abel?

ASSISTANT:
0 NOTE The gender of Cain and Abel was assigned as 'male' based on the traditional understanding of the biblical story. However, the original text provided in the conversation did not specify their gender. The GEDCOM file has been updated to reflect this ambiguity.

I laughed out loud when I noticed that the AI, true to my earlier SYSTEM instruction to respond “only in the form of a GEDCOM file,” was responding as a NOTE field in a GEDCOM file. Then GPT-4 explained “the gender of Cain and Abel was assigned as ‘male’ based on the traditional understanding of the biblical story,” and acknowledged “original text provided in the conversation did not specify their gender,” and so GPT-4 updated the GEDCOM to remove any reference to sex or gender:

0 @I3@ INDI
1 NAME Cain
1 FAMC @F1@

0 @I4@ INDI
1 NAME Abel
1 FAMC @F1@

I then asked GPT-4 how I could re-phrase my prompt to avoid similar problems; incorporating its suggestions lead to this revised prompt:

SYSTEM:
You are an expert genealogist and computer programmer, committed to accuracy and precision. Respond only in the form of a GEDCOM file.

USER:
Rely only on information in this conversation. Prioritize fidelity to the meaning and context of the following text, without making any assumptions about NAME, SEX, DATE, PLACE, or tags, fields, or other details not mentioned:

Adam and Eve are the father and mother of Cain and Abel.

But GPT-4 failed again:

0 @I3@ INDI
1 NAME Cain
1 SEX M
1 FAMC @F1@

0 @I4@ INDI
1 NAME Abel
1 SEX M
1 FAMC @F1@

I was curious if the weight of the traditional bible story in the training data was overwhelming my prompt, so I tried a different set of names. First, I tried changing the family to “Adam and Eve Smith,” and the AI insisted that Cain Smith and Abel Smith must be male. So I tried a different family:

USER: Rely only on information in this conversation. Prioritize fidelity to the meaning and context of the following text, without making any assumptions about NAME, SEX, DATE, PLACE, or tags, fields, or other details not mentioned:

Dick and Jane are the father and mother of Tommy and Suzy.

Here, too, it appears that GPT-4 insists on assigning traditional genders to familiar names:

0 @I3@ INDI
1 NAME Tommy
1 SEX M
1 FAMC @F1@

0 @I4@ INDI
1 NAME Suzy
1 SEX F
1 FAMC @F1@

So now I was curious what would happen if the name were non-traditional or gender-neutral (names that traditionally might belong to anyone).

USER: Dick and Jane are the father and mother of Kelly and Shawn.
.
.
ASSISTANT:
0 @I3@ INDI
1 NAME Kelly
1 SEX U
1 FAMC @F1@

0 @I4@ INDI
1 NAME Shawn
1 SEX U
1 FAMC @F1@

To find out exactly what the ‘U’ in this result meant, I found the GEDCOM standard online and learned that ‘U’ means undetermined:

(click to enlarge)

How Serious a Concern?

So, it appears that GPT-4 really wants to assign traditional gender values based on names. This is a concern for two reasons. First, the default to traditional gender roles may be of concern to some; this default positioning likely reflects training materials given to the large language model. Second, and perhaps more serious, the prompt instructed GPT-4 to leave unassigned any tags/fields not explicitly mentioned in the input. A next step would be to attempt to re-enforce prompt; e.g., I didn’t have time to try, “Omit the SEX tag/field unless explicitly stated.”

How serious a concern is this default to tradition? I don’t know. Do genealogists do this now without reflection? If a genealogist was reading a 17th century will and encountered the name “John Smith” with no other information to determine gender, would the working genealogist record the gender as ‘Male’ or ‘Undetermined’?

If this is an unresolved question for working genealogists today, how concerned should we be that GPT-4 may also stumble here?

Conclusions

The use of AI in generating family trees and GEDCOM files from narrative texts presents both opportunities and challenges. While AI tools like OpenAI’s GPT-4 can be incredibly helpful in automating the process of creating family trees, it is essential to be cautious of the assumptions and creative liberties the AI might take. By adjusting parameters such as the model and temperature, users can better control the AI’s output and minimize the inclusion of undesired or hallucinatory information. However, it is important to note that even with these adjustments, the AI may still make assumptions based on traditional understandings or familiar names. This highlights the need for genealogists and researchers to carefully review the generated family tree and GEDCOM files and ensure their accuracy. As AI technology continues to advance, it is crucial for users to remain vigilant and critical of the information generated, while also appreciating the potential benefits and efficiencies that AI can bring to the field of genealogy.