Datasets & code — available on request*

Sentence-level language identification

  • Source: Facebook — Assamese, Bengali, Karbi, Boro, Hindi, English
  • Source: YouTube — Assamese, Bengali, Hindi, English

Word-level language identification

  • Source: Facebook — Assamese, Bengali, Hindi, English

Chart-type classification

Jennil Thiyam, Sanasam Ranbir Singh, and Prabin K. Bora. 2021. Chart classification: an empirical comparative study of different learning models. In Proceedings of the Twelfth Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP '21). ACM, New York, NY, USA, Article 32, 1–9.

  • A chart dataset of 28 classes
  • Chart-type classification using traditional classification methods
  • Chart-type classification using deep learning-based methods
  • Chart-type classification using two attention mechanisms
  • Feature visualization using GradCAM

Chart-type studies (attention and triplet-loss based)

Thiyam, J., Singh, S.R. & Bora, P.K. Effect of attention and triplet loss on chart classification: a study on noisy charts and confusing chart pairs. J Intell Inf Syst (2022).

  • Euclidean distance calculation of confusing chart class pairs
  • Finding hard triplets from confusing chart class pairs
  • Integration of attention mechanisms into Xception
  • Triplet loss training

Manipuri OCR

  • Tesseract-based Manipuri OCR
  • Fine-tuning procedure for the existing Tesseract OCR
  • Python script for semi-supervised training to populate text corpora
  • OCR evaluation tool (provided by Tesseract)

Document segmentation

  • A sample dataset for document segmentation
  • Two segmentation models
  • Two classification models
  • Document region (textual, equational, graphic) segmentation

IndiSentiment140

  • A parallel corpus translating Sentiment140 into 22 Indian languages supported by Google Translate
  • Languages: Assamese, Bengali, Bhojpuri, Dogri, Gujarati, Hindi, Kannada, Konkani, Maithili, Malayalam, Marathi, Meiteilon (Manipuri), Mizo, Nepali, Odia, Punjabi, Sanskrit, Sindhi, Sinhala, Tamil, Telugu, Urdu
  • Same sentiment labels retained across every translated language

*Request datasets by writing to osi.iitg@gmail.com.