AI & ML interests

None defined yet.

Recent Activity

Organization Card

About Revolution Crossroads

Revolution Crossroads is a collaboration between the Smithsonian Institution and the Library of Congress, exploring how artificial intelligence can help researchers discover connections across American Revolutionary-era collections (1770–1810). The project brings together museum collections, historic newspapers, and publicly available Revolutionary War pension records.

Here on Hugging Face, we share datasets and processing outputs to support research, experimentation, and reuse. These resources reflect work in progress: we continue to refine and expand them throughout the project. Individual dataset cards document processing methods, updates, and known limitations.

Visit the Revolution Crossroads website for more information, stories, videos, and project updates.

Datasets

Our datasets bring together collection records, images, and text from three sources, along with AI-generated text and named entity extraction results. Each dataset card explains its contents, sources, processing methods, and known limitations.

Smithsonian Revolutionary Era Collections

Catalog records and object records from four Smithsonian museums: the National Museum of American History, National Postal Museum, Smithsonian American Art Museum, and National Portrait Gallery. These collections reflect many facets of life during the Revolutionary era, including letters, portraits, coins, garments, military equipment, and household items.

Chronicling America Newspapers, 1770–1810

Historic newspapers from the Library of Congress offering contemporary accounts of events, people, commerce, and daily life during the Revolutionary era and early republic.

  • Page-level dataset : Individual newspaper pages with publication metadata, links to digitized pages, original OCR text, and AI-extracted text where available.
  • Issue-level dataset : Newspaper pages grouped into complete issues, with companion PDFs, publication metadata, and text organized by page.

Revolutionary War Pension Files

Publicly available National Archives records documenting veterans’ applications for pensions, with supporting accounts from families and witnesses. These materials provide perspectives on military service, family life, and community ties during and after the war.

  • Page-level dataset : Individual pages with metadata, links to digitized images, machine-extracted text, and Citizen Archivist transcriptions where available.
  • File-level dataset : Pages grouped by pension application file, with companion PDFs, metadata, and extracted text and transcriptions organized by page.

AI Processing Results

  • Revolution Crossroads OCR Results : AI-extracted text from materials across the three collections, available by text block, page, and source file. Block-level results include layout information. This dataset supports analysis and the development of tools for finding and examining passages in context.
  • Revolution Crossroads Named Entity Occurrences : Machine-extracted mentions of people, places, organizations, dates, events, and vessels, with information linking each occurrence to its source. These results provide a foundation for ongoing work to identify entities and explore connections across the collections. Repeated mentions are retained and do not represent counts of unique entities.

Note: AI processing results may contain errors or omissions. Named entity results are also based on AI-extracted text, which may introduce additional errors. These experimental outputs should be checked against the source materials.

Buckets

Buckets hold supporting files and processing outputs that complement our datasets, making additional materials available for inspection, evaluation, and reuse. We will add public buckets as more project resources become available.

  • Revolution Crossroads VLM OCR Outputs : Raw outputs from Chandra 2 OCR processing of collection files. These include extracted text in multiple formats, structured JSON, page layout and bounding box information, extracted images, and model-generated quality assessments. Files are organized by collection and source document.

    For many uses, the Revolution Crossroads OCR Results dataset will be the preferred starting point. It combines text extracted from these outputs into a common tabular structure across the three collections, with results available by text block, page, and source file. This bucket provides the fuller outputs for users who need additional processing detail.

Spaces

Experimental tools and visualizations may be shared as Hugging Face Spaces to explore the project datasets and processing results. These are works in progress; their availability and functionality may change as the project develops.

models 0

None public yet