peter-cooper-letter-blue-1140x540

Tooling Up to Transcribe Handwritten Text with The Cooper Union

Peter Cooper is a name recognized in many spheres; he was a prominent New Yorker, a notable inventor, and a politician characterized by earnestness and a passion for social engagement. He was incredibly impactful to the development of New York and American industrialization throughout the 19th century, and a wealth of information is archived by The Cooper Union for the Advancement of Science and Art among personal letters and business correspondence.

In 2025, a grant supplied by the Metropolitan New York Library Council (METRO) made it possible for three linear feet of the The Erksine Hewitt Collection of Cooper-Hewitt papers to be digitized and transcribed. Manual transcription is a costly and time-consuming process, so Backstage proposed an alternative: would the library consider partnering on a pilot to evaluate the effectiveness of handwritten transcription performed by AI?

While the discussion around LLMs continues to require thoughtful consideration, the results are in—and the pilot went very well.

Piloting New Tools

Jamie Smith, Senior Programming Manager at Backstage Library Works, was responsible for evaluating the transcription workflow and deciding on a quality assurance protocol for the files. Backstage digitized roughly 5,000 images to TIFF, then created WEBP derivatives for internal use which are smaller in size and easier to process. These files were sent, alongside a prompt, to two different Language Learning Models (LLMs) for transcription.

The two exports were evaluated by a third LLM that was employed to score the results and determine which was the more accurate for a given image. The model was also able to pull unstructured metadata from the transcriptions and provide inventories that describe the sender, addressee, and notable dates, names, places, and organizations from each page of correspondence. Validity of the structured metadata elements was especially crucial to the library as it would aid in completing research requests, so Backstage standardized these keywords through partial human review and batch editing.

Cooper1
Cooper2
Cooper3

There are copyright and personal privacy considerations when utilizing LLMs. However, the collection is old enough and the content such that neither factor would be a concern—and for The Cooper Union, this made all the difference in determining if the collection would be a good fit for the pilot. Additionally, while the data was sent to external parties through a custom API, the files themselves could only be accessed through authorization protocols tied to a unique account. “So, while the system was open, access to the content was limited to our profile alone,” explains Jamie.

Customization & Considerations

Like many projects at Backstage, we performed a lot of testing and samples to make sure the results would be exactly as the library expected. To start with, a total of 8 different models were evaluated according to their accuracy, and only the top performing platforms were selected to be used in processing. Beyond the technology considerations came very practical ones; one aspect that needed standardization was how certain content, like crossed out sections, would be returned. “In the end, we opted for bracketed information that would help the reader better understand the orientation of the content,” confirms Calista Donohoe, Digital Collections & Services Librarian at The Cooper Union. If a note was added into the margin, for example, a bracketed note would be added beside the text explaining its irregular position on the page.

While the transcription was good, it was not without its limitations. Reference requests sometimes need a few iterations to hit on the perfect result, but given the cost and timeline (only a week of processing versus months or even years of manual transcription), The Cooper Union is excited about the level of discoverability.

What are the lessons learned? In future projects, mapping text outputs to their place inside the archival image would make a great quality of life upgrade for researching. And, while the transcription process went quickly when compared to analogue processing, Jamie reported that time was lost in repeat requests or repeat processing during off hours. Jimmy Wolff, Digitization Operations Manager with Backstage, summarizes the mutual feelings about this project perfectly: “Being able to explore the workflow with a solid programming team and a library that shared our feelings of curiosity and caution made this process really rewarding.” 

Accessible Results

Cursive literacy has become increasingly uncommon in the 21st century, and even those who are able to read it may struggle with penmanship that is unique or downright flawed. Archival collections, too, are simply not as discoverable, and not as concurrently accessible as digitized facsimiles.

“We called it Unlocking Peter Cooper’s New York because this project truly unlocked the potential of these resources,” says Mary Mann, Cooper Union’s College Archivist. Since the project’s conclusion earlier this year, the files have been integrated into a digital repository and The Cooper Union was able to host an exhibit kicking off access to the collection.

To learn more about Backstage’s approach to AI in future projects, reach out to us at info@bslw.com.

Learn More About

Share this post

Looking for Something?

Search our site below