I could not be more excited about sharing this announcement and learning with you. Registration opens Tuesday 20 February 2024, 1 PM ET, for “AI Genealogy Seminars: From Basics to Breakthroughs,” one of eleven virtual courses offered this summer by the National Genealogical Society’s GRIP Genealogical Institute (formerly the “Genealogical Research Institute of Pittsburgh”). I will be teaching ten sessions, from basics to breakthroughs in AI Genealogy, and I am humbled to be joined by seven distinguished colleagues who were students (though “fellow pioneers” would be more accurate, as I learned as much from them as they did from me) in my Level 1 and/or Level 2 NGS AI Genealogy courses last fall and this winter. More information is included below, and at the websites noted below.
Course: AI Genealogy Seminars: From Basics to Breakthroughs Coordinator: Steve Little Date: 23-28 June 2024 Venue: GRIP Virtual Session Registration Opens: 1 PM ET, Tue 20 Feb 2024
Description:
As the AI Program Director at the National Genealogical Society (“NGS”) and a pioneer in the field, Steve Little will navigate participants through the foundational concepts to the frontiers of AI Genealogy. His sessions will chart the evolution of AI Genealogy, from its early stages to predictive trends in 2024. Sessions will cover practical skills in prompt engineering, bot building and GPT customization, and the latest in AI Genealogy advancements. Steve’s comprehensive expertise will provide attendees with the tools to not only grasp AI basics but also to apply sophisticated AI strategies to their genealogical research. The AI Genealogy Seminars offer a unique opportunity to learn from the first-hand experiences of industry leaders during the initial year of large language model integration into genealogy. Their insights will shape the course content and provide a diverse perspective on the evolution of this field. Featured experts include Blaine Bettinger, Maureen Taylor, Judy Russell, Dana Leeds, Nicole Dyer, Mary Kircher Roddy, and Mark Thompson.
NGS online course “Empowering Genealogists with Artificial Intelligence – Level One” or commensurate experience (approx. 20+ hours of hands-on AI genealogy experience). The GRIP 2024 course “AI Genealogy Seminars” is NOT for a genealogist’s first experience with large language models or other AI tools; however, if a student has 20+ hands-on hours with ChatGPT Plus (GPT-4) when the course begins, then an updated refresher of the basics and intermediate aspects of AI Genealogy will ensure the student is up-to-speed to take advantage of the advanced topics and special content instructors.
About NGS’s GRIP Genealogy Institute: Formerly the “Genealogical Research Institute of Pittsburgh”, GRIP 2024 is “the” event for genealogists and family historians who want to develop their skills while meeting new friends in a collegial and collaborative community.
Open GeneaGPT (the community-built genealogy AI tool) has been updated from version 2 to version 3 as of Monday 22 January 2024. This is a significant update, transforming Open GeneaGPT into a smarter and more engaging companion for exploring family history, making it easier and more enjoyable for everyone. It brings changes like better conversations, more helpful suggestions for next steps, ensuring a friendly and insightful journey into the past.
OpenAI’s custom GPTs are amazing tools for creating specialized bots that can perform various tasks, such as extracting data; describing documents, images, records; generating reports, stories, images; and much, much more. These GPTs are tools for the way you work; you build the AI tool you need, or find one on the store shelf. I have been exploring genealogy prompt engineering over the past year, building genealogy GPTs since November, and I have learned a lot about how to create and use them effectively. Recently, OpenAI launched their GPT Store, where users can share their custom GPTs with other ChatGPT Plus subscribers. This is a great opportunity to discover new bots and learn from others. I have shared five or six of my custom genealogy GPTs in the store, which are related to family history research (more are being finished in the lab). Custom GPTs are available to ChatGPT Plus subscribers ($20/month) for no additional cost in the ChatGPT Store.
These are the kinds of genealogy AI tools that I teach students to create; no programming skills are necessary. ChatGPT can interview you, to ask you what kind of AI tool you’d like to create. Then I can show you how to fine-tune the tool to suit your exact genealogy workflow. It won’t fetch you a Snickers Bar or promise to solve a 120-year family history mystery, but within the realm of what large language models can do today, this one does okay.
If you’d like hands-on, step-by-step instruction on how to build your own custom GPTs and other specialized genealogy AI tools, your own flock of bots, the National Genealogical Society is now enrolling for Prompt Engineering and Specialized AI Tools for Genealogists, which starts at the end of January 2024; six hours of instruction over four weeks is just the start; also includes a collaborative study group, sharing successes and learning from failures. The first section of 50 seats sold-out in days, so a second section has been opened. Learn more here: https://www.ngsgenealogy.org/ai/
Jump straight the the growing list of Genealogy Bots here or access them at OpenAI Store by searching for bots with the term "genealogy." Custom GPTs are free for ChatGPT Plus subscribers ($20/month). I also teach genealogists and educators how to make their own custom genealogy GPTs, hands-on, step-by-step; enrolling now.
Friends, you may have heard the announcement that the OpenAI directory of custom GPTs is now being unrolled to ChatGPT Plus users. Custom GPTs represent an advancement in AI usefulness, marking a step towards more customizable and versatile AI tools. According to the OpenAI, these custom GPTs are described as a means to “create for a specific purpose,” highlighting their adaptability and user-centric design. Wharton professor Ethan Mollick in his article “Almost an Agent: What GPTs can do,” emphasizes the current state and future potential of these tools, noting, “GPTs show a near future where AIs can really start to act as agents.” This statement underscores the transitional nature of GPTs as a bridge between current AI capabilities and the more autonomous agents of the future. Simon Willison, in “Exploring GPTs: ChatGPT in a trench coat?” offers a practical perspective, stating, “The combination of features they provide can add up to some very interesting results.” His experience reflects the innovative possibilities that arise when various capabilities of GPTs are combined. Together, these insights from OpenAI, Mollick, and Willison paint a picture of GPTs as transformative tools, offering both a glimpse into the future of AI agents and a practical platform for current applications.
I’ve got four GPTs (now five) that I’m publicly testing (these are genealogy-related GPTs or genealogy-adjacent; another half-dozen others are still in the lab). I think of these GPTs as little AI tools that you can create, save, repeatedly re-use, and share. Custom GPTs, also referred to as bots, assistants, or agents (though these terms aren’t technically synonymous), represent a method to save a bundle of prompts, custom instructions, and abilities (image analysis, image generation, document reading, etc.) in a profile that you can use and share. For us as genealogists, this means that when we find a prompt, series of prompts, or set of custom instructions, to reliably accomplish a genealogically useful task, we can save that process as one of these GPTs; then, when we need to accomplish that task again, that tool, that bot, that GPT, is already in our AI toolbox.
Months ago, professional genealogist Yvette Hoitink created and shared the first genealogy GPT (to my knowledge), Dutch Genealogy Bot, available to ChatGPT Plus users at https://chat.openai.com/g/g-MMm3v0QX3-dutch-genealogy-bot. When asked for a short summary of its abilities, the bot replied: “As the Dutch Genealogy Bot, I specialize in guiding you through the process of researching Dutch ancestry, using resources and insights exclusively from Yvette Hoitink’s Dutch Genealogy website. I can provide detailed information and methodologies for tracing Dutch heritage, and direct you to specific articles on DutchGenealogy.nl for further guidance and authentic information.”
These four five little bots are my initial efforts, for example:
Genealogy Eyes Look at images, photos, and documents through the eyes of a family historian. Try it from your phone! Take a snapshot of a cemetery headstone, document/record, or anything else, and, using the official ChatGPT app, upload the image, say a little about the image and what you want, and click Send. https://chat.openai.com/g/g-gmIAn5mh6-seer-of-roots
Lingua Maven A linguistic expert, I combine dictionary precision, usage panel insights, and style guide expertise. My skills encompass a vast lexicon, dynamic thesaurus, and in-depth knowledge of language evolution, etymology, and dialects. I am a comprehensive resource for analysis and interpretation. https://chat.openai.com/g/g-Rxyt3Xww1-lingua-maven
Genealogy Summarizer Create useful summaries from texts, images, documents, photos, records, and more. This bot looks at what you give it, determines (as best it can) what it is, and suggests several ways to summarize the item. You can then ask follow questions about the item, or collect all the suggested summaries. https://chat.openai.com/g/g-Kg79HuRVD-genealogy-summarizer
Sam the Digital Archivist Kinda like a spicy librarian. Open GeneaGPT’s over-caffeinated genealogy and family history friend. Embark on a journey through your past with our customized genealogist bot, designed to delve into your ancestry and lineage. Discover your roots while having fun and learning genealogical methods. https://chat.openai.com/g/g-v6WgbVnba-sam-the-digital-archivist
PS: Learning how to make these custom GPTs is a significant portion of the focus of the Empowering Genealogists series class, Level 2: Prompt Engineering and Specialized AI Tools for Genealogists, which starts at the end of January 2024: https://www.ngsgenealogy.org/ai/
Identifying the people in a photo or letter collection is a top priority for family archive projects. It’s incredibly frustrating to have a beautiful stack of old photos or precious family letters and not know who is in them.
Unfortunately, identifying people in old collections is easier said than done: last names may be omitted, nicknames could be used instead of formal names, and letters might lack dates.
In larger collections, this problem is reduced because people are mentioned many times by different people. Each time they are mentioned by someone new, a new clue may be found. Each mention provides a clue that may help identify them.
However, in large collections, comparing and analyzing these clues can be overwhelmingly complex. Could ChatGPT simplify this challenge?
Comparing Items
Comparing the information found in each item in a collection is a daunting task. Comparing items in a large collection is especially challenging and requires both art and science to succeed.
On the science side, a complete research log is key. I use a spreadsheet-based research log, tuned to the needs of each project, to capture facts and dates about each item. This makes it relatively easy to find relationships between items.
Here is a simple example to illustrate the spreadsheet I use:
On the art side, a good memory for names, facts, and patterns is a huge bonus. Like many genealogists, I live for the “Eureka! I’ve seen that name before!” moment after a multi-hour research session.
As much as I enjoy the adrenaline rush of these eureka moments, I don’t like to rely on my memory to make them happen. Developing processes that make these patterns easier to find is crucial to successful research.
Clustering
Clustering is a technique for finding the relationships between groups of people or things. Even if they don’t use the term, genealogists use clusters every time they research. For example:
The phone book clusters people who live in the same town into groups of people with similarly spelled last names.
Land maps cluster people into groups who own land that is close together.
The Leeds Method clusters people into groups who descend from the same grandparents.
Census records cluster people into groups of people who lived in the same building at the same time.
Friends, associates, and, neighbours (FAN) clusters group people together who are familiars of our family members.
Finding and understanding these clusters, and the relationships between the people in the cluster, can generate clues that are critical to breaking down genealogy brick walls!
The Manual Approach to Clustering
Before automating anything, it is important to learn how to do it the “manual” way.
To make it easier to identify clusters of people I start with a detailed research log. I track the names of each person mentioned in each item in the log. Then, using advanced spreadsheet techniques, identify the people who are mentioned together in the same items.
Knowing that these people are likely connected, I can then focus my efforts on determining how they are related. They might be friends, coworkers, or family. Knowing how they are related makes it easier to figure out who they are.
For large collections with hundreds of items and thousands of people, I use a network analysis tool named Gephi, to create diagrams for deeper analysis.
While invaluable, clustering with Excel and Gephi can be daunting initially. Even with a meticulous research log, this analysis technique can require 50 to 100 hours of practice to use with confidence.
How To Use ChatGPT To Simplify Clustering
Getting Started: Prompt Planning In a Nutshell
When approaching a new problem with ChatGPT, consider the following steps:
Determine how to provide the AI with the necessary information for the task.
Decide the best role for the AI to assume when tackling the task.
Define the specific task(s) you wish the AI to accomplish with the provided information.
Choose the format in which you want the AI to deliver its findings.
Prompt Planning Example: Clustering People Mentioned in a Collection of Letters
I’ll walk you through the prompt planning process that I used with a real-world family archive project so that you can try this, or something similar, yourself. Note: The data mentioned in this article is sourced from an actual client project, with permission granted for its use here.
How to Provide the AI With the Information
As the archival information is stored in an Excel spreadsheet, the best option would be to find a way to use the existing spreadsheet directly in ChatGPT.
ChatGPT Plus added a feature today that was previously part of the “Advanced Data Analysis” plugin. This proved especially good timing for this example because ChatGPT Plus can now read an Excel spreadsheet directly. As such, there was no need to convert the spreadsheet to text, or a CSV file, to place it in the chat prompt.
Assign a Role to ChatGPT
Next, it is important to assign a role for the AI to adopt when performing the task. The role assignment is critical to do first as it sets the context for how additional instructions will be understood by ChatGPT.
As I was trying to perform a relatively complex data analysis task with archival documentation, I wanted ChatGPT to assume a role that specialized in this type of work. This would give it the context to understand the rest of the prompt.
Prompt: Please act in the role of an expert data analyst who has a deep knowledge of archive and records management.
Define the Task you Want To Do
After some testing, it became clear that there were several steps in this particular task that worked best when described separately.
Find the Information
First, ChatGPT needed clear guidance on how to find the information in the spreadsheet.
Prompt: Look for the column titled “People Mentioned” on the “Letters” tab in the provided Excel file.
Understand the Information
Next, it needed guidance on what kind of information it was going to work with. Ensuring that ChatGPT understood what it was working on, and not just where it was found, would make it possible to perform tasks on the information using natural language.
Prompt: This column contains a comma separated list of the people mentioned in a letter. Each row in the spreadsheet refers to a different letter.
Extract The Information
Once I had “taught” the AI how to find and understand the people mentioned in the spreadsheet it was time to get down to the real work. The goal was to identify the people who were mentioned most frequently in the letters and then to find the cluster of people who were also mentioned with them in the letters.
As this collection of letters is between family members, I hoped that the resulting clusters of people would know each other somehow and that the clusters would help identify them.
Prompt: Identify the five people who are mentioned most frequently across all the letters. Then, identify the five people who are most frequently mentioned in the same letter as each of the people you just identified.
Choose the format for the AI to Present Its Results
While there are a few common forms for a cluster analysis to take, the first approach I tried was in a table.
Prompt: The response should be a table, with the first column containing the most frequently occurring names, and the second column containing a comma-delimited list of the names which are most frequently found along with them.
The Final Prompt
After testing each step in the prompt plan, here is the final prompt.
Prompt: Please act in the role of an expert data analyst who has a deep knowledge of archive and records management.
Look for the “Letters” tab in the provided Excel file. Each row in this spreadsheet contains information about a letter. The column “People Mentioned” on this tab contains a comma separated list of the people mentioned in a letter. Each row in the spreadsheet refers to a different letter.
Identify the five people who are most frequently mentioned across all the letters. Then, identify the five people who are most frequently mentioned in the same letter as each of the people you just identified.
The response you provide should be a table, with the first column containing the most frequently occurring people, and the second column containing a comma-delimited list of the names which are most frequently found in the same letter with them.
The Final Prompt – Cluster Diagram Edition
Although a table format helps see who else is mentioned in the same letters as the most frequently mentioned people, there is a graphical format that provides additional insight called a Cluster Diagram.
Cluster diagrams show relationships not just between individual people, but between groups of related people. This additional layer can provide insights that are easily missed in a table.
This follow-up prompt was also modified slightly to present a larger cluster of people for analysis. However, a filter was used to limit the number of people in the chart so that it was easier to read. Simply said, I wanted to identify the people who had an ongoing relationship with the most frequently mentioned people.
Prompt: Please draw me a cluster chart that shows relationships between the people mentioned in the same letter as those you identified in the table. Draw a connecting line between two individuals if they are mentioned together in at least two letters. The line weight should be the same for all relationships.
Why This Cluster Diagram Is So Amazing!
It’s Useful For Research
The diagram clearly shows two clusters of people that are mentioned in the same letters as the most frequently mentioned people. And, the two clusters are connected by one individual, Elvira Caroline Ingersoll.
There is no need to go deeply into the family used in this example, but I will say that the cluster diagram is spot on! Elvira was a prolific writer and recipient of letters. In one generation, she was a daughter who communicated frequently with her siblings and parents. In the next generation, she was a mother who communicated extensively with her husband and children.
Any family historian or archivist would assume that this was a possibility if they saw this cluster chart. In this case, this clue would quickly put them on the useful track!
It’s Conceptually Easy to Create
It is difficult to understand the concept of a cluster and how to determine which information is beneficial to include within it. ChatGPT was able to understand my direction in natural language and retrieve the necessary information directly from a spreadsheet.
It’s Technically Easy to Create
Filtering a cluster down to the most useful items is technically difficult to do in Excel or Gephi. The fact that I got an easy-to-read and easy-to-understand chart so easily was nothing short of mind-blowing. This shows huge potential for “quick and dirty” cluster analysis, something that is relatively difficult for genealogists and archivists to do today.
Total Time Saved Using ChatGPT
The way I usually make these charts takes about 30 minutes of spreadsheet work and 60-90 minutes of work in Gephi. Keep in mind, this is after about 100 hours of learning and practice.
To create these charts in ChatGPT took only 30 minutes. Although there was some conceptual overlap with learning how to do this manually, there were none of the technical challenges associated with using complex, and occasionally buggy, pieces of software.
To be fair, Gephi can handle more information and more complex types of clustering and visualization than ChatGPT. While I will still need to use Excel and Gephi for more complex analysis, I won’t need Gephi for this type of quick and dirty cluster analysis anymore.
Navigating Challenges
Test Each Step Separately
This prompt involved several steps, and each step needed to be tested separately before it could be combined with the other steps. It was difficult to see where things were going wrong when I tried to test several steps at the same time.
New Chat Windows
I found it best to frequently start a new chat window during testing. Hallucinations and unexpected results were introduced when I tested several steps in the same chat window, particularly if something went wrong in a previous step.
The Importance of Clean Data
Starting with a clean, well-structured, research log is key to the success of any data analysis project. The old axiom of, “garbage in, garbage out” should never be forgotten.
Clustering with ChatGPT is about Insight, Not Tools
With ChatGPT, clustering becomes less about juggling complex software and more about spotting the helpful patterns in your data. Yet, there’s no silver bullet here; it still requires a knack for seeing the patterns in the first place.
Conclusion and Final Thoughts
When I started this article, I had high expectations for some parts and very low expectations for others. In the end, my expectations were exceeded in all areas!
Clustering is a very useful technique for identifying people in a family archive and determining the relationships between them. Although, it can be technically challenging to undertake this type of analysis.
ChatGPT had no difficulty accessing and understanding archival information in a complex Excel spreadsheet. Because of that, I had no difficulty interacting with the information in the spreadsheet using natural language.
ChatGPT made quick work of creating cluster tables and diagrams that showed the people most frequently mentioned in the archive. It also had no difficulty identifying the people who were mentioned in the same letters along with them.
Amazingly, ChatGPT could apply filtering rules to these clusters of people, a conceptually and technically difficult activity to do with traditional clustering tools. The ease with which ChatGPT produced clear, comprehensible cluster diagrams, a feat usually only possible through significant learning and practice, cannot be understated.
I am super excited about the potential of ChatGPT to analyze archives. Dozens of use cases immediately come to mind where it will greatly speed up my work.
My next ChatGPT clustering challenge though will come from the field of genetic genealogy … DNA Match clusters!
I’m very excited to hear what you think about using ChatGPT to perform cluster analysis. What do you want to learn from Clustering? What mysteries are waiting to be found in your family archive?
A common goal when working on family archive projects is to figure out who the people are that are included in photographs or mentioned in letters. Identifying them can be crucial to answering questions about your family history and can lead to new clues to follow up on.
Although, as anyone who has tried to do this knows, it can be a painstaking and frustrating process. Letters rarely include the information needed to identify everyone. As letters are usually between people who know everyone they’re writing about, they tend not to include last names, or even worse, first names.
I’ve developed Excel-based approaches over the years for identifying groups of people in these situations. Although, they can be time-consuming to put together, and require specialty skills in Excel to use them.
It occurred to me that I might be able to use artificial intelligence to identify people more easily. This blog post will describe how I tested this idea and what I learned that could help you do the same in your own genealogy research.
Before diving into the fictitious example I used to test this idea, it is helpful to understand how this kind of research is done “manually.”
How to Manually Identify People in a Family Archive
While there are many approaches for identifying people mentioned in a letter, one of the most common techniques used by genealogists is to figure out how the different people mentioned in the letter may be related. They might be friends, co-workers, or family. If you can figure out how they are connected, it is easier to figure out who they are.
After all, it’s easier to find several needles tied together in a haystack than it is to find an individual needle in the haystack.
How to Identify Family Members
For the rest of this article, the names that I use will all be taken from this example, fictitious, family tree.
When I suspect that people mentioned in a letter could be family members, I compare the names of the people to their family tree and try to find close relatives with those names. The assumption is that people tend to write about their immediate family.
For example, let’s say a letter written by “Clara” includes a line that says, “Walter and I went down to the train station to pick up Joe.” While the family tree might include dozens of people named Joe, Walter, and Clara; it will have fewer families (and hopefully, only one) where they are all immediate family members.
While this approach makes intuitive sense, it isn’t easy to use in practice. The problem is that online trees don’t have a button labeled, “Show me all of the families that have a person named Joe, Walter, and Clara in them.” As such, this approach requires repetitive searches of a family tree looking for clues about which family group might be the right one. Alternatively, spreadsheets, or third-party tools designed for this kind of search, can also be used.
Given the challenges in doing these searches the manual way, I decided to see if this problem could be solved more easily using artificial intelligence tools.
Using ChatGPT to Search a Family Tree
Given that ChatGPT is particularly good at finding patterns, it seemed a natural fit for this type of search. Now the big question was how to get ChatGPT to search a family tree? I needed a plan.
Prompt Planning with ChatGPT
Whenever I am working out how to approach a problem using ChatGPT, I think about the following:
How will I provide the AI with the information needed to do the work?
Which role should the AI take on when performing the work?
What is the work that I want the AI to do with the information I provide it?
What is the format that I want the AI to present its findings in?
I’ll walk through my planning process step by step so that you can try this, or something similar, yourself.
How to Give ChatGPT Your Family Tree
The first, and most difficult challenge in this example was to get ChatGPT the information from the family tree. To do the kind of search that I’m interested in, ChatGPT needed to be able to see the people in the tree and understand their relationship to each other.
Thankfully, there is a family tree format that ChatGPT can understand.
The GEDCOM File Format
Image Generated by DALL-E 3
The Genealogical Data Communications format, or GEDCOM for short, was created by the Church of Jesus Christ of Latter-day Saints as a way for exchanging family tree information between computer programs. The information that can be transferred using this format includes information about the people in the tree, (like name, and birth, marriage, and death dates) as well as information about the relationships between the people (like parent, child, and family group).
Of all of the formats for family tree information in use today, you may wonder why GEDCOM is good to use with ChatGPT?
Why GEDCOM Works Well with ChatGPT
Besides the fact that GEDCOM files contain the information needed for this search, there are a few things about the GEDCOM format that make it well-suited to working with ChatGPT.
GEDCOM is a text-based format and ChatGPT excels at working with text.
Even though ChatGPT was trained on information created several years ago, the version of GEDCOM used by the major genealogy companies is several years old. This means that the information that ChatGPT has in its training data is still accurate today.
The GEDCOM format, which has been in use for 40 years, has been the subject of thousands of online articles. In fact, Google found over 4 million web pages that mention GEDCOM! As a result, ChatGPT has been well-trained in how the GEDCOM format works.
Now that we know that the GEDCOM format is a good way to provide information to ChatGPT, how do we get our family tree into the GEDCOM format?
How to Export an Ancestry Family Tree in GEDCOM Format
To generate a GEDCOM file of my family tree at Ancestry, I completed the following steps. Starting within Ancestry’s Tree View:
Select the three dots menu in the left navigation bar.
Select the “Tree Settings” menu.
Click the “Export tree” link.
Click the “Download your GEDCOM file” button.
Note, that depending on the size of your tree, it can take some time to create the export file before it can be downloaded.
I then saved the GEDCOM file to a known location on my computer, so that I could use it in the next steps.
Now that my family tree was exported to GEDCOM format, it was time to build the ChatGPT prompt.
Assigning a Role to ChatGPT
The process of building a prompt, often referred to as prompt engineering, starts by assigning a role for the AI to adopt when performing the work. The role assignment is important to do first because it sets the context for how additional instructions will be understood by the AI.
To understand why role assignment is important, consider how you ask different people to do work for you. For example, the way that you would ask your 12-year-old child to clean up your yard would be different than the way you would ask a person who works for a professional yard cleaning service. Even though you have very similar goals for them, you would phrase the request for each of them in a very different way because of who they are.
As I was trying to form a complex genealogy search of a GEDCOM file, I wanted ChatGPT to assume a role that specialized in this type of work:
PROMPT: Please act in the role of a professional genealogist who has a deep understanding of the GEDCOM file format.
While the work portion of the prompt seems incomplete by itself, combined with the role assignment in the first step, it had the context to be understood by ChatGPT.
Finally, I needed to tell ChatGPT how I wanted the information it found to be presented to me.
How Should the AI Present Its Findings?
In this example, my goal was to look at the results for clues about family groups. So, I wanted the response to include information about the relationships between the people found. And, because I was going to do these searches frequently, I wanted the results to be easy to interpret at a glance.
PROMPT: Should you find a family group that you believe includes these people, create a table that lists the full name of the people in the family group in one column, and their relationship to Joe in another column.
The Complete Prompt
After testing several different approaches, this is the final prompt. Note, that I’ve shared some of the failed attempts at the end of the article in the “Challenges to be Aware of” section.
PROMPT: Please act in the role of a professional genealogist who has a deep understanding of the GEDCOM file format.
I would like you to analyze the file that I will provide to you next. Please search the file for family groups that include the names Joe, Walter, and Clara.
Should you find a family group that you believe includes these people, create a table that lists the full name of the people in the family group in one column, and their relationship to Joe in another column.
After submitting the prompt, I opened the GEDCOM file in a text editor so that it would be easy to copy the file to provide it to ChatGPT. In my case, I used Notepad++, but you could do this with Notepad, Wordpad, or any other text editor.
Once I had selected and copied all of the text, I pasted the text directly into ChatGPT’s prompt box and then clicked the submit button.
I was very happy to see that it found the correct family group, and displayed them in a way that was easy to confirm and check for additional clues!
Other Complex Searches Tested
I tried several different complex searches that I regularly come up against when doing this kind of project.
Finding More Than One Family Group
In a real-world search, it is likely that I would find more than one family group that included the names that I was looking for.
PROMPT: Please act in the role of a professional genealogist who has a deep understanding of the GEDCOM file format.
Please search for family groups that include the name Terry.
Should you find a family group that you believe includes this person, create a table that lists all of the people in the family group in one column, and their relationship to that person in the other column. Should you find more than one family group, create an additional table for each additional family group.
Focus on People That Were Alive at The Time
One way to zero in on the right family group is to only include people who were alive at the time the letter was written.
PROMPT: Please act in the role of a professional genealogist who has a deep understanding of the GEDCOM file format.
Please search for family groups that include a person named Terry who was alive in 1960.
Should you find a family group that you believe includes this person, create a table that lists all of the people in the family group in one column, and their relationship to that person in the other column. Should you find more than one family group, create an additional table for each additional family group.
This response is particularly interesting as it shows that ChatGPT, acting in the role of a genealogist, knows how to interpret a request for living people.
This is a textbook example of a natural language search.
And many, many more…
I tried several other, increasingly complex searches, and they all worked as long as the information that I was searching for was included in the family tree.
Challenges to Be Aware Of
ChatGPT Plus Can Only Accept 25,000 Characters
The most important limitation to be aware of is that there is a limit on how much text you can ask ChatGPT to process. Because I used ChatGPT Plus in my testing, the limit is about 25,000 characters.
When I tried this test with a sample tree with hundreds of people in hundreds of family groups, ChatGPT said it was “too big.” When I performed my test on a smaller family tree with only a few dozen family groups, it worked successfully. As a result, I consider the approach used in this article a good proof of concept for me, and others, to use as a starting point for searches of larger and more family trees.
There is also a less well-understood limit that is based on the “complexity” of the file being processed. In the context of a GEDCOM, I believe this comes into play when there are more types of facts, and more relationships between people to track. Although, I wasn’t able to find a clear description of this limitation.
I expect that as ChatGPT evolves, these limitations will decrease.
GEDCOM Files Might Contain Personal Information
Be careful when exporting your family tree GEDCOM file format. If your tree contains private information, so will your GEDCOM file.
Ancestry’s GEDCOM export utility exports all the facts in your tree. If you would like to export a portion of your family tree, or only certain facts from your family tree, you will need to use a tool that supports this.
Family Tree Maker, for example, supports the partial export of a family tree.
Exercise Caution with Follow-up Searches
My testing was most successful when I used a fresh chat session in ChatGPT. When I tried follow-up searches in the same chat session, errors in the responses went up dramatically. When I spotted mistakes, I asked ChatGPT to double-check its results and explain how it came to its conclusion. In every case, it found the correct answer on the second try and apologized for its mistake.
Because of this issue, I quickly learned to follow up each response with a “please double check and explain your results” prompt.
Based on my testing, I believe that there are two likely sources for these errors:
The errors might come from answers generated in previous prompts. In other words, ChatGPT mixed its previous responses up with my subsequent requests.
The amount of information in the session grew too long after multiple requests, so information was being “forgotten” because of the space taken up by previous questions.
Final Thoughts
ChatGPT Plus can read and search GEDCOM formatted family trees and correctly interpret the genealogical information in them.
It can also do complex searches of family trees using natural language. These complex queries can include searches for multiple people, family groups, relationships between people, and the time or place that people lived.
False responses were generated by some of the tests, especially when multiple follow up questions were used in the same chat session.
Like all genealogy research, results from ChatGPT need to be treated as clues that require further investigation before they can be relied upon as fact.
At the time of this writing, the approach used in this article is limited to “small” trees. Alternative approaches, or improvements to ChatGPT Plus, will be necessary to search larger trees.
I’d Love To Hear From You
Have you tried any alternative approaches for complex searches of a family tree?
Do you know of a way to search larger family trees with a different approach? If so, please let me know in the comments below.
Artificial intelligence (AI) excels at teaching basic concepts, serving as a personalized tutor across various fields, including genealogy. In this way, AI can act as our tutor whether we are trying to learn terms and skills related to genealogy or to another part of our life.
Recently, when prompted, it did a nice job of explaining the Genealogical Proof Standards (GPS). I went a step further, requesting an example of the third step: analysis and correlation of information.
Image generated by DALL-E 3 with title corrections using Procreate
I appreciate how AI tools like ChatGPT provide quick responses to our specific questions. Searching for an example of genealogical analysis and correlation online would be time-consuming. But with AI, it’s quick!
Another benefit is that we can ask AI to make its response simpler or more advanced. For instance, I can request ChatGPT to simplify the explanation for a sixth grader’s understanding or to elaborate as if I were a college freshman. In this way, we can quickly generate, and even modify, the explanations and examples we need.
However, we must be cautious about AI’s potential inaccuracies, often referred to as “hallucinations.” ChatGPT is getting better in this regards, but we need to be cautious. AI tools like ChatGPT tend to provide less accurate responses to highly specific questions than to more general information and explanation requests.
The remaining section highlights ChatGPT’s tutoring skills as it provides a detailed example of genealogical analysis and correlation.
So, in the words of ChatGPT, here is an example of analysis and correlation of genealogical materials:
ChatGPT
Certainly! Analysis and correlation in genealogical research involve critically examining all the collected information, then comparing and synthesizing this data to draw well-founded conclusions about ancestors and their relationships. Here’s an example of how this might work:
Scenario:
Imagine you’re researching an ancestor named John Smith who lived in Virginia in the early 1800s. You’ve collected several records that mention individuals named John Smith, and now you need to analyze and correlate this information to ensure you’re actually tracing the right person.
Image generated by DALL-E 3
Records Collected:
A 1820 U.S. Federal Census record lists a John Smith in Albemarle County, Virginia, with a household of six.
An 1823 marriage record from Albemarle County for a John Smith and Elizabeth Johnson.
An 1830 U.S. Federal Census record lists a John Smith in Augusta County, Virginia, with a household of four.
An 1832 land purchase deed in Albemarle County for a John Smith.
An 1840 U.S. Federal Census record lists a John Smith in Albemarle County with a similar age and family composition to the 1820 census.
Analysis and Correlation Process:
Census Records: You begin by examining the age categories of the individuals in the households of the 1820 and 1840 Albemarle County census records. They match up with the expected ages of John Smith, his wife, and children, allowing you to infer they likely represent the same family. However, the 1830 census in Augusta County appears to be for a different John Smith, as the family size and location don’t match.
Marriage Record: The 1823 marriage record provides a wife’s name, which can be checked against the ages of the female members of the household in the 1840 census to see if it matches the expected age of Elizabeth Johnson Smith. If it does, this supports that you have the correct family unit.
Land Records: The 1832 land deed in Albemarle County suggests that John Smith either remained in the county or returned after 1830. You would correlate this with tax records, which might show a continuous presence in Albemarle County, suggesting the John Smith in the 1830 Augusta County census may indeed be a different person.
Consistency Across Documents: You also notice that the ages of the children across the census years seem consistent, suggesting a continuous and consistent growth of the family, further supporting the conclusion that these records pertain to the same John Smith.
Neighbors and Associates: Looking at the names of neighbors in the census records, you might find that certain families consistently appear near John Smith, which supports that you’re tracking the same individual across the censuses.
Naming Patterns: If you have access to probate or church records, you might find that John named his children after his siblings or parents, a common practice, and this can be corroborated with earlier records of John Smith’s supposed parents and siblings.
Image generated by DALL-E 3
Conclusion:
After carefully comparing the information, you conclude that the John Smith in the 1820 and 1840 Albemarle County census records and the 1832 land deed is the same person, and he is the individual you’re researching. The 1830 Augusta County record is deemed to be a different John Smith. You have accounted for the whereabouts of your John Smith consistently from 1820 to 1840, and you have a marriage that fits within the timeline and location expected for your ancestor.
This example demonstrates how genealogists use analysis and correlation to build a case for an ancestor’s identity and life events, ensuring that the person they are adding to their family tree is indeed the correct individual.
Your Turn
Have you tried using AI as a tutor? What did you ask it and how did it help? I’d love to hear from you about how you’re using AI alongside genealogy.
If you’re curious about how artificial intelligence (AI) might help you become more efficient, consider experimenting with citations. The large language models (LLMs) that have become popular in the last year might prove useful in this area. I’ve been using ChatGPT to experiment with AI in the last month as a student in the National Genealogical Society (NGS) course, Empowering Genealogists with Artificial Intelligence taught by Steve Little. Steve has a great blog here: https://aigenealogyinsights.com/. In the NGS course, we had four sessions about the four main LLM transformations: summarize, extract, generate, and translate. Today I experimented with having ChatGPT help me generate citations.
Book Citation
When learning how to create citations, it’s a good idea to start with a basic book citation. The principles for citing books are carried over to citing other publications in genealogy research, like websites. For genealogy citations, we use the Evidence Explained (EE) style. This style is based on the Chicago Manual’s Humanities Style (as opposed to its author-date style). The format for a book citation is the same in both EE and Chicago.
I asked ChatGPT 4 to help me cite a book in Chicago format. It then provided me both the format for a bibliographic entry and the footnote/endnote. Here’s the full response:
To cite the book “Research Like a Pro: A Genealogist’s Guide” by Diana Elder and Nicole Dyer in Chicago format, you would generally use the following structure:
Bibliography Entry: Elder, Diana, and Nicole Dyer. Research Like a Pro: A Genealogist’s Guide. [Place of Publication]: [Publisher], [Year of Publication].
Footnote/Endnote:
Diana Elder and Nicole Dyer, Research Like a Pro: A Genealogist’s Guide ([Place of Publication]: [Publisher], [Year of Publication]), [Page number].
In these citations, replace [Place of Publication], [Publisher], [Year of Publication], and [Page number] with the appropriate information from the book. The page number is only necessary in the footnote or endnote and should correspond to the specific part of the book you are referencing.
It’s important to note that the title of the book is italicized in both the bibliography and the notes. Additionally, in the notes, only the first author’s last name is listed first, followed by the first name, while the second author’s name is in the standard format.
Next I asked it how to create a shortened version of the reference note for use as a subsequent note. Here’s the response:
For a shortened version of the reference note for “Research Like a Pro: A Genealogist’s Guide” by Diana Elder and Nicole Dyer, used in subsequent notes in Chicago format, you would generally include just the primary author’s last name, a shortened version of the title, and the page number being referenced. It would look like this:
Elder and Dyer, Research Like a Pro, [Page number].
For instance, if you are referencing page 45 in a subsequent note, it would be:
Elder and Dyer, Research Like a Pro, 45.
This shortened format is used after the first full citation of the work in your footnotes or endnotes to make subsequent references more concise.
ChatGPT did a great job with providing the accurate format for a Chicago book citation. The Chicago Manual of Style, 17th ed., gives the same format on p. 799, section 14.101.
Census Citation
I asked ChatGPT to create a census citation. I started with a citation that I had already made for myself for my ancestor, Eliza Ann Isenhour. I separated the citation I had previously made into multiple lines, removed the punctuation, added an extra word (webpage near accessed) and put the civil division into the same line as the county and state – just to see if it knew where to put that. Here’s my prompt:
You are an expert genealogist. Make a citation for Eliza Isenhour on the 1850 census following the Evidence Explained style manual used by genealogists. Here is the information about Eliza’s enumeration:
1850 U.S. census Polk, Montgomery County, Arkansas population schedule p. 405a dwelling 6 family 6 Eliza Ann Isenhour in Barnet Isenhour household database with images Ancestry https://www.ancestry.com/discoveryui-content/view/9812411:8054 webpage accessed 17 November 2023 citing NARA microfilm publication M432, roll 28
The response:
To create a citation for Eliza Isenhour in the 1850 U.S. Census following the Evidence Explained style, you would structure it like this:
“1850 U.S. census, Montgomery County, Arkansas, Polk, population schedule, p. 405a, dwelling 6, family 6, Eliza Ann Isenhour in Barnet Isenhour household; database with images, Ancestry (https://www.ancestry.com/discoveryui-content/view/9812411:8054 : accessed 17 November 2023); citing NARA microfilm publication M432, roll 28.”
This citation follows the general format recommended by the Evidence Explained style guide, which includes the year and type of record, geographic location, specific details of where the record is found (page, dwelling, family), the individual’s name and household context, the database and its URL, the date of access, and the specific archival reference (NARA microfilm publication and roll number).
I was pretty impressed! It got almost everything right. This citation is usable. It did not place Polk, the civil division, in the correct position according to Evidence Explained, 3rd ed., p. 237. The civil division should follow the words “population schedule.” I prompted ChatGPT to put Polk where it belongs, and it did. I then asked it to make a few other tweaks, like removing the a after 405a, which is the page number given in the source information from Ancestry, and asked it to add the word (stamped) after 405. Then I asked it to italicize the publication website, Ancestry.
It’s important to check for accuracy while ChatGPT is still evolving. Because LLMs aren’t search engines, they sometimes produce incorrect information, according to what it thinks you want to hear. When I asked ChatGPT to make a citation for Eliza Ann Isenhour on the 1850 census, without providing the details from her actual household, it suggested that I follow this format:
“1850 United States Federal Census, [State], [County], [Township or City], enumeration district [ED], page [Page Number], dwelling [Dwelling Number], family [Family Number], Eliza Isenhour; digital image, [Name of Database or Collection] ([URL of the database] : accessed [Date of Access]); citing National Archives and Records Administration microfilm [Microfilm Number].”
This is an okay suggestion, and I love that it created a template that you can fill in. However, it incorrectly tells you to include the enumeration district (ED). EDs were not used until the 1880 census. Read the full conversation here: https://chat.openai.com/share/b7638e29-d6f9-431f-85a9-d2a008ebeb7d
Why do LLMs produce incorrect suggestions? The GPT in ChatGPT stands for Generative Pre-trained Transformer. The AI has been trained by reading a lot of content and is best suited for helping you with creative tasks or helping you summarize a block of text that you provide. The way it works is by predicting what usually comes next when talking about specific topics. So, if you provide ChatGPT with the way you like your citations to look, then give it some information to use, it can probably help with that. While we’re still learning the capabilities of ChatGPT and other LLMs, it’s important to check their work. This will help you know if you can use it for the tasks that will help make your work more efficient.
Shortening Citations
Another usage you might try for ChatGPT and citations is asking it to shorten a citation you’ve already made, when you want to use it for a subsequent note. I returned to my first conversation with ChatGPT about creating a citation for Eliza Isenhour where I provided the details. After the full reference note citation was tweaked the way I like it, I asked it to shorten the citation for a subsequent note. Here’s the response:
For a subsequent note, a short form citation for Eliza Isenhour in the 1850 census can be created as follows:
“1850 U.S. census, Montgomery Co., Ark., Polk, p. 405 (stamped), Eliza Ann Isenhour.”
In this short form citation, the details are condensed, providing just enough information to identify the specific record without repeating all the information from the full citation. It includes the year, the abbreviated state and county, the civil division, page number, and the name of the individual.
My response:
Please include population schedule, abbreviated, and the dwelling and family numbers, which can also be abbreviated.
ChatGPT’s response:
Certainly! Here’s the revised short form citation including the requested elements:
“1850 U.S. census, Montgomery Co., Ark., pop. sch., Polk, p. 405 (stamped), dwl. 6, fam. 6, Eliza Ann Isenhour.”
In this version, “population schedule” is abbreviated as “pop. sch.,” and “dwelling” and “family” are abbreviated as “dwl.” and “fam.,” respectively. This format maintains the necessary details while being concise for a subsequent note.
This can be a starting point for helping you figure out how you want to shorten citations for subsequent notes in a particular piece of genealogical writing that you’re creating.
Conclusion
It seems possible that LLMs could help us start drafting our citations. However, without me providing the details for the 1850 census that I already knew would be helpful, it had a much harder time. Maybe we can train our chatbots to give us citations in the format we prefer after providing them with additional examples and training.
The book citation was great, so using ChatGPT for help citing common source types, like books and articles online, is a great usage right now.
Your Turn
If you experiment with asking ChatGPT to create a citation for you, share below in the comments how it turned out. Was the citation formatted the way you’d expect according to the style manual?
Note: If you are sharing any private information as you practice making citations, turn off your chat history so it’s not seen by the engineers training the chatbot. See https://help.openai.com/en/articles/7730893-data-controls-faq for more info.
More AI and Genealogy Resources
Bettinger, Blaine.”Unlocking Family Secrets with AI | Findmypast,” 22 March 2023. YouTube. https://www.youtube.com/watch?v=exepLKC72Ts. This webinar is a great introduction to LLMs and Artificial Intelligence for genealogy and covers the purpose of LLMs, what they can do, and what they can’t do.
——. “10 ChatGPT Prompts Every Genealogist Needs to Know | Findmypast,” 10 May 2023. YouTube. https://www.youtube.com/watch?v=EbRXzd2SmNM.This gives some great ideas for how to use ChatGPT in genealogy.
Little, Steve. “Empowering Genealogists with Artificial Intelligence 6 September 2023.” YouTube. https://www.youtube.com/watch?v=npQaRJbzE1s. This is an overview of AI for genealogy and some of the capabilities of ChatGPT.
As genealogists, we find ourselves at the intersection of history and technology. Today’s example comes from the evolving field of AI-generated imagery. Earlier today, Steve Little from AI Genealogy Insights highlighted a resource that addresses a common challenge with DALL-E 3: fine-tuning the generated images to fit our specific needs by using seeds.
As part of a private Facebook group for Steve’s NGS “Empowering Genealogists with Artificial Intelligence” course, a Twitter post by Rowan Cheung was shared which sheds light on the concept of using ‘seeds’ to refine these images. While I had come across the term earlier in the week, it wasn’t until this demonstration that the methodology clicked for me—and it’s a game changer!
What’s the problem?
This and all following images generated using DALL-E 3
Often, when you try to make small changes to a generated image, it instead generates an entirely new image. And that can be frustrating!
For example, above is an image of a black cat with a sign that says “Trick or Treat” (with DALL-E 3’s notorious spelling mistakes).
I liked the cat and the scene, but I wanted to change the sign. So I prompted “Generate another image similar to #1 but have the sign say ‘Boo!’”
The new image changed a lot more than just the sign. It includes a different, younger black cat and different background though it’s similar.
Using seeds can help us to produce additional images that are much similar to the original image.
What is a “Seed”?
To better understand the concept, I turned to ChatGPT for an explanation of what “seeds” mean in the context of AI. Here’s an analogy from the first part of its response:
“Alright, imagine you have a magic book that gives you a random page number every time you ask it. But sometimes, you want to get the same ‘random’ page number every time you ask, so you can show your friend the cool picture on that page. That’s kind of what a seed does in AI.”
So, if we tell DALL-E 3 we want to modify a certain seed or page number, it does a pretty good job of giving back a similar image with the modifications we asked for.
Using Seeds
After reading through the Tweet mentioned earlier, I started playing with seeds (again). This time, I understood the process a lot better. And it worked!
My first success was a baby seal. I will use a series of prompts to get the desired image doing what is called “prompt chaining.”
Prompt #1: Draw an adorable seal with large eyes
Of the two seals DALL-E 3 generated, this was the “adorable baby” seal I chose to work with.
Prompt #2: What’s the seed for image 1?
Since the baby seal I wanted was the first of the two generated images, I asked for the seed for image 1. It responed “1122301494.” Remember, this is like me now knowing what “page” DALL-E 3 has the image on so I can go back and modify that page!
Prompt #3: Modify the image with seed 1122301494: add a beach scene with a starfish -ar 7:4
When I’m asking DALL-E 3 to modify a seed, I start with the phrase “modify the image with seed [x].” Next, I can ask it to add, remove, or edit something. And finally, I can ask for specific aspect ratios (-ar): 1:1 for square, 7:4 for wide, or 4:7 for tall. (I have also learned I can just use the words square, wide, or tall!) I love changing a square image into “wide” or “tall” which actually expands the scene!
We now have the SAME ADORABLE BABY SEAL with a wider beach scene and a starfish. WOW!!!
And just to try it again…
Prompt #4: Modify the image with seed 1122301494: add his mother -ar 7:4
In hindsite, I don’t think I needed to specifiy the aspect ratio since it was already wide. But I’m still learning!
And once again we have the SAME ADORABLE BABY SEAL with his mom!
Breathing Life Into Family History
And now a bit more from ChatGPT based on my input:
Transforming images with seeds isn’t just about tweaking a picture until it’s perfect; it’s a gateway to visual storytelling that can vividly illustrate our family histories. Whether we aim to elevate a photograph from ‘like’ to ‘love’, or we wish to infuse static images with dynamic action, the possibilities are endless.
Take, for instance, the journey I embarked on with a single image: a photograph of a young Confederate soldier. The original image captured a moment, but I envisioned more. I wanted to broaden the narrative. So, I expanded the image—widening its scope to include additional figures, thereby crafting a richer tableau.
There was also the matter of authenticity. Through prompt chaining, the soldier’s uniform had faded to brown in the “photograph.” With careful editing, I restored the uniform’s gray hue, maintaining historical accuracy while breathing new life into the image.
Join me as I continue to explore the potential of using seeds in AI to create more accurate and personalized visual stories.