NDLI: Unconstrained handwritten document retrieval

Content Provider	SpringerLink
Author	Cao, Huaigu Govindaraju, Venu Bhardwaj, Anurag
Copyright Year	2010
Abstract	With the ever-increasing growth of the World Wide Web, there is an urgent need for an efficient information retrieval system that can search and retrieve handwritten documents when presented with user queries. However, unconstrained handwriting recognition remains a challenging task with inadequate performance thus proving to be a major hurdle in providing robust search experience in handwritten documents. In this paper, we describe our recent research with focus on information retrieval from noisy text derived from imperfect handwriting recognizers. First, we describe a novel term frequency estimation technique incorporating the word segmentation information inside the retrieval framework to improve the overall system performance. Second, we outline a taxonomy of different techniques used for addressing the noisy text retrieval task. The first method uses a novel bootstrapping mechanism to refine the OCR’ed text and uses the cleaned text for retrieval. The second method uses the uncorrected or raw OCR’ed text but modifies the standard vector space model for handling noisy text issues. The third method employs robust image features to index the documents instead of using noisy OCR’ed text. We describe these techniques in detail and also discuss their performance measures using standard IR evaluation metrics.
Starting Page	145
Ending Page	157
Page Count	13
File Format	PDF
ISSN	14332833
Journal	International Journal of Document Analysis and Recognition (IJDAR)
Volume Number	14
Issue Number	2
e-ISSN	14332825
Language	English
Publisher	Springer-Verlag
Publisher Date	2010-11-16
Publisher Place	Berlin, Heidelberg
Access Restriction	Subscribed
Subject Keyword	Image Processing and Computer Vision Pattern Recognition
Content Type	Text
Resource Type	Article
Subject	Computer Vision and Pattern Recognition Software Computer Science Applications

Sl.	Authority	Responsibilities	Communication Details
1	Ministry of Education (GoI), Department of Higher Education	Sanctioning Authority	https://www.education.gov.in/ict-initiatives
2	Indian Institute of Technology Kharagpur	Host Institute of the Project: The host institute of the project is responsible for providing infrastructure support and hosting the project	https://www.iitkgp.ac.in
3	National Digital Library of India Office, Indian Institute of Technology Kharagpur	The administrative and infrastructural headquarters of the project	Dr. B. Sutradhar bsutra@ndl.gov.in
4	Project PI / Joint PI	Principal Investigator and Joint Principal Investigators of the project	Dr. B. Sutradhar bsutra@ndl.gov.in Prof. Saswat Chakrabarti will be added soon
5	Website/Portal (Helpdesk)	Queries regarding NDLI and its services	support@ndl.gov.in
6	Contents and Copyright Issues	Queries related to content curation and copyright issues	content@ndl.gov.in
7	National Digital Libarray of India Club (NDLI Club)	Queries related to NDLI Club formation, support, user awareness program, seminar/symposium, collaboration, social media, promotion, and outreach	clubsupport@ndl.gov.in
8	Digital Preservation Centre (DPC)	Assistance with digitizing and archiving copyright-free printed books	dpc@ndl.gov.in
9	IDR Setup or Support	Queries related to establishment and support of Institutional Digital Repository (IDR) and IDR workshops	idr@ndl.gov.in

CMATERdb1: a database of unconstrained handwritten Bangla and Bangla–English mixed script document image

A character image restoration method for unconstrained handwritten Chinese character recognition

ICDAR2009 handwriting segmentation contest

Using topic models for OCR correction

A graph-based approach for segmenting touching lines in historical handwritten documents

A categorization system for handwritten documents

Special issue on document recognition and retrieval 2009

A Bayesian-based method of unconstrained handwritten offline Chinese text line recognition

Large scale document image retrieval by automatic word annotation

Unconstrained handwritten document retrieval

Similar Documents

CMATERdb1: a database of unconstrained handwritten Bangla and Bangla–English mixed script document image

A character image restoration method for unconstrained handwritten Chinese character recognition

ICDAR2009 handwriting segmentation contest

Using topic models for OCR correction

A graph-based approach for segmenting touching lines in historical handwritten documents

A categorization system for handwritten documents

Special issue on document recognition and retrieval 2009

A Bayesian-based method of unconstrained handwritten offline Chinese text line recognition

Large scale document image retrieval by automatic word annotation

Unconstrained handwritten document retrieval