Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

The sources were the GLOBALISE transcriptions of the Dutch East India Company archive (4.35 million pages, from the 1600s to the 1790s), digitized Dutch and American newspapers, and some ship logbooks. Boston engineer Jesse Waites built a model cascade to process them. Typesafe AI's Jev screened the text first. Claude Haiku then read the few dozen passages Jev flagged, translated them and extracted dates and places. A Claude Code agent then compared each transcription against the scan of the handwritten original. A Deep Research assistant picked the targets, returning thirteen candidates ranked partly on whether an answer could be checked against an original page. The project builds on historian Benjamin Breen, who searched the same transcriptions and found a 1615 ship's journal recording dodos. Founders should take the architecture from this. A cheap model filters, a mid-size model extracts, and an agent checks every claim against the primary source. Waites estimated that reading the Company pages by hand would take about 70 years. That gap is the market for archival search tools, but only if their outputs can be traced back to the original documents.