In progress
Multimodal Image-to-Music Generation
This project asks a deceptively simple question: if a person can look at a photograph and
hum something that fits it, what would a model have to learn in order to do the same? The
work sits at the join of computer vision and audio generation, and treats music as a
second description of the same scene rather than a decorative accompaniment to it.
The approach is encoder–decoder. A pretrained vision backbone maps an image into a dense
representation; that representation is projected into a shared embedding space and used to
condition an audio generation decoder. The interesting engineering is in the middle — the
projection — because that is where the model either learns a real correspondence between
visual and musical structure or collapses into producing the same pleasant loop for every
input.
Alongside the model, the team is assembling an aligned image–music dataset and defining how
this task should even be scored. Objective audio metrics say very little about whether a
piece suits a picture, so evaluation combines retrieval-style measures with
structured human listening comparisons.
Research questions
- Which visual attributes — palette and colour temperature, composition, density, subject
matter — carry information the decoder can actually use?
- Do those attributes map onto musical ones in a stable way: tempo, mode, instrumentation,
dynamic range?
- How much of the correspondence is genuinely learned, and how much is the model falling
back on a prior over "generic pleasant music"?
- What evaluation protocol separates a musically competent output from an appropriate
one?
- Status
- In progress
- Started
- 10 August 2026 — ongoing
- Modalities
- Image → Audio
- Areas
- Multimodal AI · Generative Models
- Researchers
- Gurpreet Singh · Lamia Qamar
In progress
Text-to-Comic Generation
Generating a single illustration from a sentence is close to solved. Generating twenty
illustrations that tell one coherent story is not. This project takes written narrative as
input and produces sequential comic pages — and the hard part is everything that has to
stay true from one panel to the next.
The pipeline runs in stages. A script is segmented into narrative beats, one per panel;
a layout planner decides how those beats sit on a page and in what reading order; an
image generator renders each panel; and dialogue is extracted and placed into balloons
positioned so the eye reaches them in the right sequence. Each stage is a research problem
on its own, but the failures compound — a good panel in the wrong place still breaks the
page.
The dominant open problem is consistency. A character must remain recognisably the same
person across panels drawn independently, in changing poses, framing and lighting. The
team is working on conditioning strategies that carry identity through the sequence, and
on an evaluation that scores narrative faithfulness and character consistency separately
from raw image quality.
The work builds directly on the lab's earlier storytelling and cross-modal generation
projects, which handled narrative structure and modality translation independently.
Research questions
- How should continuous prose be segmented into discrete panels — by event, by dialogue
turn, or by a learned boundary?
- What conditioning keeps a character identifiable across independently generated panels?
- Can panel layout and reading order be planned by the model rather than templated?
- How do you measure whether a generated page tells the story it was given?
- Status
- In progress
- Started
- Ongoing
- Modalities
- Text → Image sequence
- Areas
- Generative Models · Multimodal AI
- Researchers
- Trina Banerjee · Mukhthikka · Gurpreet Singh
In progress
Multimodal AI in Gene Editing
Gene editing generates evidence in several forms at once — nucleotide sequence, microscopy
and assay imaging, and a large published literature describing what has already been tried.
Researchers routinely hold all three in mind together. Most computational tools handle one
at a time.
This project applies the lab's cross-modal methods to that gap: encoding sequence, image and
text evidence into a shared representation so a model can reason across them the way a
biologist does. It is the same architectural question as our image-to-music and
text-to-comic work — how do you build a space where two very different signals become
comparable — pointed at a domain where being right matters a great deal more.
The near-term aim is modest and concrete: support target selection and outcome prediction,
and be explicit about the confidence attached to each. The project sits squarely in the
lab's AI-for-social-good area, alongside the earlier drug–drug interaction work.
Research questions
- Which representation lets sequence, image and literature evidence be compared without
one modality dominating the others?
- Does adding imaging and literature actually improve target selection over
sequence-only baselines, or only appear to?
- How should a model in this domain express uncertainty so that a biologist can act on
it?
- What does responsible evaluation look like when the downstream application is
biological rather than aesthetic?
- Status
- In progress
- Started
- Ongoing
- Modalities
- Sequence + Image + Text
- Areas
- Multimodal AI · AI for Social Good
- Researchers
- Sudipta Patil · Gurpreet Singh
In progress
International Research Collaboration
Eight of the lab's members already work from outside India, and this project makes that
distribution the subject rather than a logistical fact. It pairs the lab's AI methods with
research in international business, ESG and media — the disciplines its members
actually came from — and runs the work across time zones from the outset.
The substantive question is what AI changes about cross-border business and communication:
how organisations report and are held to account, how markets and audiences behave across
cultures, and where automated analysis genuinely helps rather than flattening the
differences that matter. The methodological question is quieter but just as real —
what a distributed research team needs in order to produce something coherent.
Scope and outputs are being defined with the collaborating partners; this entry will be
expanded as they are fixed.
Research questions
- What can AI-driven analysis tell us about international business and ESG reporting
that conventional methods miss?
- How do media narratives and consumer behaviour differ across markets, and can a model
capture that without erasing it?
- Which parts of a cross-border research programme benefit from automation, and which
depend on a person who knows the context?
- What working practice lets a distributed team sustain one line of enquiry?
- Status
- In progress
- Started
- Ongoing
- Focus
- International business · ESG · Media
- Areas
- AI Ethics & Bias · AI for Social Good
- Researchers
- Ritika Kanojia · Alimpia Roy · Gurpreet Singh