Multimodal AI
Vision, text and audio in one system — captioning, visual question answering, cross-modal alignment.
About
CAAI is an experimental research lab for students entering artificial intelligence — a place to learn the craft by running real projects, not by reading about them.
01 — Mission
CAAI Lab supports students entering the field of artificial intelligence, machine learning, computer vision and cross-modal systems. We provide a guided research environment where beginners learn research methodologies, implementation practices, paper writing and experimentation under structured mentorship.
Members take a question and carry it end to end — literature, method, implementation, evaluation, write-up. The work that comes out is theirs, and it goes out under their names.
02 — Organization
CAAI Lab operates under The Korean Academy, a registered organization in India. The lab supports students learning Korean and preparing for academic or professional transitions to Korea or other countries, while building research profiles for international higher education opportunities.
03 — How the lab works
04 — What we work on
Vision, text and audio in one system — captioning, visual question answering, cross-modal alignment.
Fairness, accountability and transparency — measuring where models fail people.
Language and diffusion models for text, image, audio and code generation.
Healthcare, education and environmental problems where better models matter.
Join us
The lab takes graduate research assistants and undergraduate research interns each term. If you can program and you're willing to be patient with an experiment, we want to hear from you.