Language transition on temple inscriptions using Tesseract OCR and T5 model
Lakshmi Revathi Krosuri, Sai Sri Harsha ChintalapudiTranslating ancient languages like Sanskrit poses significant challenges in Natural Language Processing (NLP) due to its complex morphology, low-resource status, and script variations. This study proposes a hybrid Optical Character Recognition (OCR) and NLP-based translation system to enhance Sanskrit-to-English translation accuracy. The system integrates Tesseract OCR for high-precision text extraction and SpaCy for linguistic preprocessing, including tokenization, part-of-speech tagging, and named entity recognition (NER). To address Sanskrit’s grammatical complexity and contextual dependencies, a fine-tuned T5 Transformer model is employed for context-aware translation, text correction, and summarization. The framework is evaluated on a diverse Sanskrit corpus , demonstrating significant improvements in word recognition accuracy, BLEU scores, and overall processing efficiency compared to existing OCR-NLP models. By combining OCR with deep learning-based NLP techniques, this approach enhances translation quality while contributing to the digital preservation and accessibility of Sanskrit literature, supporting further research in computational linguistics and historical text analysis.