KanVQA: Visual Question Answering in Kannada
An interactive research demonstration for multimodal visual question answering in the Kannada language. Upload any image, type a question in Kannada, and view predictions across baseline and state-of-the-art architectures.
🏆 SOTA PaliGemma 3B (56.20%)
⚡ ResNet-50 + mBERT Baseline (35.30%)
📊 20,000 Train / 1,000 Test Pairs
🏛️ Manipal Institute of Technology
💡 Quick Question Suggestions (ತ್ವರಿತ ಸಲಹೆಗಳು):
📸 ➜ 🧠
Ready for Prediction
Upload an image and question, then click 'Predict Answer'
📊 Class Probability & Prediction Analysis
Probability distribution will appear here
🖼️ Preset Sample Gallery (Click to load image & question)
📈 Benchmark Accuracy Across 6 Modeling Paradigms
Evaluated on the 1,000-example human-verified test dataset.
| Model Architecture | Paradigm | Exact Match Acc | BLEU-1 | ROUGE-L | CIDEr | Trainable Params |
|---|---|---|---|---|---|---|
| 🥇 PaliGemma-3B (QLoRA) | Fine-tuned VLM (0.385% params) | 56.20% | 53.92 | 57.33 | 147.16 | 11.30M / 2.93B |
| 🥈 ResNet-50 + XLM-RoBERTa | CNN–Transformer Fusion | 39.20% | 38.10 | 40.50 | 82.40 | 303M (Full) |
| 🥉 ResNet-50 + mBERT Baseline (.pth) | CNN–Transformer Fusion | 35.30% | 34.60 | 37.20 | 75.10 | 142M (Full) |
| 4️⃣ BLIP-VQA-base (Fine-tuned) | Fine-tuned Generative VQA | 19.30% | 18.50 | 19.80 | 35.60 | 385M (Full) |
| 5️⃣ ViLT-B32 (Zero-shot) | Zero-shot Multimodal | 15.30% | 14.10 | 15.60 | 22.10 | 0 (Zero-shot) |
| 6️⃣ BLIP-VQA-base (Zero-shot) | Zero-shot Generative VQA | 9.50% | 8.90 | 9.20 | 12.40 | 0 (Zero-shot) |
| 7️⃣ CLIP ViT-B/32 (Zero-shot Ranking) | Zero-shot Contrastive | 3.20% | 3.00 | 3.10 | 5.20 | 0 (Zero-shot) |
📚 KanVQA Dataset Breakdown
Photographic images from Microsoft COCO paired with parallel English and Kannada visual questions and answers.
| Dataset Split | # QA Pairs | # Unique Images | Quality Assurance Status |
|---|---|---|---|
| Training Split | 20,000 | 16,924 | Machine-translated (IndicTrans2 + Google API) |
| Evaluation Split | 1,000 | 982 | Human-Verified & Text Corrected by Native Speakers |