Resources
Demos, datasets, and code
Systems the group has put online, and datasets or code released for research use.
Demo systems
Datasets & code — available on request*
Word polarity detection using syllable features
- Language: Manipuri
Sentence-level language identification
- Source: Facebook — Assamese, Bengali, Karbi, Boro, Hindi, English
- Source: YouTube — Assamese, Bengali, Hindi, English
Word-level language identification
- Source: Facebook — Assamese, Bengali, Hindi, English
Chart-type classification
Jennil Thiyam, Sanasam Ranbir Singh, and Prabin K. Bora. 2021. Chart classification: an empirical comparative study of different learning models. In Proceedings of the Twelfth Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP '21). ACM, New York, NY, USA, Article 32, 1–9.
- A chart dataset of 28 classes
- Chart-type classification using traditional classification methods
- Chart-type classification using deep learning-based methods
- Chart-type classification using two attention mechanisms
- Feature visualization using GradCAM
Chart-type studies (attention and triplet-loss based)
Thiyam, J., Singh, S.R. & Bora, P.K. Effect of attention and triplet loss on chart classification: a study on noisy charts and confusing chart pairs. J Intell Inf Syst (2022).
- Euclidean distance calculation of confusing chart class pairs
- Finding hard triplets from confusing chart class pairs
- Integration of attention mechanisms into Xception
- Triplet loss training
Manipuri OCR
- Tesseract-based Manipuri OCR
- Fine-tuning procedure for the existing Tesseract OCR
- Python script for semi-supervised training to populate text corpora
- OCR evaluation tool (provided by Tesseract)
Document segmentation
- A sample dataset for document segmentation
- Two segmentation models
- Two classification models
- Document region (textual, equational, graphic) segmentation
IndiSentiment140
- A parallel corpus translating Sentiment140 into 22 Indian languages supported by Google Translate
- Languages: Assamese, Bengali, Bhojpuri, Dogri, Gujarati, Hindi, Kannada, Konkani, Maithili, Malayalam, Marathi, Meiteilon (Manipuri), Mizo, Nepali, Odia, Punjabi, Sanskrit, Sindhi, Sinhala, Tamil, Telugu, Urdu
- Same sentiment labels retained across every translated language
*Request datasets by writing to osi.iitg@gmail.com.