At UW-Stout Polytechnic’s University Archives and Area Research Center, thousands of handwritten records, such as genealogical surname cards, court files and naturalization documents, hold decades of regional history that are effectively locked away in cursive scripts that make the records difficult to search.
This summer, applied mathematics and computer science senior Jake See began the first phase of a project – the Automated Handwriting-Transcription Tool – to bring these materials, and the information they hold, to researchers.
“The archives is exploring how artificial intelligence and other computational tools can support the work of preserving and providing access to historical records. The handwriting recognition project is one part of a broader effort to explore how new technical tools can assist archival work. A successful handwriting recognition system could make historical records faster to process and easier to search and review,” See said.
During his internship, See applied classroom concepts and independent research toward the real-world problem of the barriers of manual transcription within archival research, and contributed to the university archives’ research and innovation goals.
His internship was a subcomponent of Emerging Technologies Archivist David Statz’s project, the Local AI Classification and Metadata Engine, which looks at the archives’ unorganized digital folders and produces natural language descriptions of each file, identifying what an image or document contains and suggesting a likely category for sorting and searching.
“The projects aim to spare archivists and researchers the slow, manual process and time of transcribing records one card at a time. By creating these tools, additionally, these materials become searchable for students, genealogists and community members who can’t visit in person,” said Statz, who contributed technical expertise to See in his project.
Automated Handwriting-Transcription Tool
Over his 200-hour internship, See established the technical groundwork for an automated handwriting-transcription tool built entirely on free, open-source software.
“One challenge facing the archives is the large amount of handwritten material in its collections. While many of these records have been digitized as images, the information within those images cannot easily be searched without first being converted into machine-readable text,” he said.
Using handwritten genealogical records from the archives as an initial collection, the project has involved two main stages.
The first stage was developing a computer vision pipeline to turn raw archival scans into usable data for handwriting recognition. Using image-processing and computer vision techniques, See developed a system that identifies and extracts individual record cards from scanned pages, locates the relevant fields within each card, distinguishes handwritten content from the preprinted form, removes unwanted printed elements, and segments the remaining handwriting into individual lines.
See developed tools for reviewing the pipeline’s output and manually adjusting crops when necessary, as well as a provenance system to keep the processed data traceable back through its source records. He also developed tools for transcribing and auditing the resulting images, producing a curated dataset of more than 1,500 handwritten lines.
See’s project focused on index cards from the 1950s, ’60s, and ’70s that recorded birth, death and marriage announcements from the Dunn County News during those decades. The initial sample of approximately 2,400 cards – a mere fraction of the total amount contained in the archives – was selected by Statz as a standardized, consistent handwritten record for the project. Head Archivist Victoria Stewart framed the project by sorting the cards to isolate and only collect the birth, death and marriage cards since those most closely align with the county records for the initial sample. After sorting, Stewart and archives student worker Renee Smith scanned the cards for See’s use.
“These cards (in their current form) live in boxes and bags in the closed stacks. Researchers would have to manually search the cards to find this information or manually search the newspaper. In working with students, I can advise on archival practices and the limitations or restrictions of access for the analog materials. The students then create and build new technologies to bring my vision of access to life,” Stewart said.
The second stage uses the prepared images to develop the handwriting recognition system itself. The project combines OpenCV, an open-source computer vision library used extensively in the image-processing pipeline, with PyLaia, an open-source handwritten text recognition framework used to train a model to convert the segmented handwriting into text. There is also a record-keeping system that traces every processed image back to its original document, as well as tools for transcribing and auditing the results.
The complete data and model-training pipeline has been validated, and the project is moving into training and evaluating its first baseline recognition model.
“Human review was built into the process from the beginning, combining the efficiency of automated processing with human verification and oversight. Because accuracy and the integrity of historical records are essential, human oversight remains an important part of the process. These tools are intended to supplement human expertise by automating repetitive work while keeping people involved in reviewing and verifying the results,” See said.
The scanning of cards also continues as more samples can be uploaded to the database. With a new CZUR scanner – an overhead, camera-based digital scanner – archives can scan and download an item in 1.5 seconds, saving countless hours of work.
See plans to share the tools as free, open-source software once complete. Because the work is built on openly available, non-commercial technology rather than subscription services, other libraries and archives, particularly smaller institutions with limited staff and budgets, could adapt it for their own handwritten collections.
“Jake’s work establishes the foundational tool on which the archives will build as the tool becomes more complex to review more elaborate, cursive scripts found in other historical documents. The open-source foundation promotes transparency, sustainability and collaboration across teaching, research and administration,” Statz added.
Collaborations creating student opportunities
See’s internship was built upon an ongoing collaboration between the archives and the university’s mathematics, statistics and computer science department, initiated by Stewart to create opportunities for students to gain hands-on technical experience while developing tools that address real needs within the archives.
Mathematics and computer science Professor Seth Dutter had advised archives staff on the overall project in the earliest stages of development and recommended See for the internship based on his coursework performance and completion, technical skill and excitement for the project. Understanding the value and importance of the internship, Library Director Roxanne Backowski provided support to hire and fully fund the internship.
“The partnership gives students an opportunity to gain substantial experience in areas such as software engineering, computer vision and machine learning while working through the challenges of developing a real system from end to end. That experience can help prepare students for technical careers or further study while allowing them to apply what they learn in the classroom to projects that directly support the archives’ work,” See said.
Statz thinks student internships play a central role in providing them with hands-on experience with real AI systems while advancing campus priorities in artificial intelligence.
“Applied learning is one of UW-Stout Polytechnic’s defining strengths. Jake’s internship provides an excellent example of how students gain hands-on experience with emerging technologies while addressing real-world challenges,” he said.
Local AI Classification and Metadata Engine
Wandering up and down the archive aisles, Statz gazed at the hundreds of aging ledgers, boxes of documents and photos, and countless other objects and thought about the archives’ digital files.
He was struck by an inspiring question: Could a digital folder of completely unlabeled images and documents with no metadata or tags – descriptive data of digital objects used to search for or categorize the objects – be turned into something structured and searchable, all without relying on commercial artificial intelligence platforms or uploading files to the cloud?
The answer is yes. Statz’s goal for the Local AI Classification and Metadata Engine is to offer a sustainable, modular open-source AI engine that adapts to any context where unstructured images or documents need interpretation and generates clear, meaningful descriptions.
“Everything happens locally, with no data leaving the machine. It’s fast, private and flexible. We created this because manual tagging takes time and rarely stays consistent. Archives, libraries and departments often have backlogs of unclassified material that can’t easily be searched or reused. This tool gives those projects a head start – a structured lens on what’s there – so teams can focus on review and refinement rather than starting from scratch,” Statz said.
For example, when fed a photo of the Old Wilson House Mansion in downtown Menomonie, the tool was able to label the image as depicting “a large Victorian-style house with a pointed roof, a chimney and a porch. The presence of a tree in front of the house suggests that the house is located in a yard. The architectural style and the time period are estimated to be from the 19th century. The house appears to be a significant landmark, possibly a historical site or a well-preserved example of Victorian-era architecture.”
The newly created label will help researchers more quickly find this historical photo using keyword and metadata searches, such as “Victorian-style house,” in the archives’ digital collection.
Statz’s first success was with a prototype completed in May 2025. It runs on three local language models and performs image, text and optical character recognition analysis through automation instructions written in Python, a computer programming language often used to build websites and software. The resulting descriptions are then processed through a local database that groups related terms and concepts, allowing the system to narrow and rank potential category matches before final selection. That version remains the foundation for all future scaling.
In June 2025, a simplified institutional version was built to demonstrate feasibility using Ollama, an open-source tool that runs AI models locally.
“Together, the two trials showed that independent local AI classification is both practical and effective. The next step is to scale the system for broader use,” Statz said. “Once the base system is fully operational, it can be expanded with automated ingestion workflows and other enhancements to improve speed, efficiency and ease of use from the end user’s perspective.”
Statz has collaborated with the BOSE research cluster at UW-Eau Claire, which provides the power to experiment with larger models and multimodal workflows.
From engineering to communication and counseling, manufacturing, marketing and design, construction, supply chain and more, UW-Stout Polytechnic is preparing graduates to meet the needs of a rapidly evolving workforce by embedding AI training in all of its degree programs. The university’s AI Innovations Committee, centers, and services collaborate with community, business and industry partners to help Wisconsin leverage AI-driven solutions that put it ahead of the curve.