Research

Projects, running and finished.

Twelve projects since 2022, across cross-modal generation, language, biology, ethics and security. Four are running in the lab right now.

01 — Currently running

Four active projects.

Three of them sit in the same territory — taking one modality and generating another — and the fourth carries the lab's methods across borders and disciplines. Members join one of these on arrival and work it through to a submission.

Image Music
Fig. 1 — Encode image, condition decoder, generate audio
Project [09] Cross-modal generation
In progress

Multimodal Image-to-Music Generation

This project asks a deceptively simple question: if a person can look at a photograph and hum something that fits it, what would a model have to learn in order to do the same? The work sits at the join of computer vision and audio generation, and treats music as a second description of the same scene rather than a decorative accompaniment to it.

The approach is encoder–decoder. A pretrained vision backbone maps an image into a dense representation; that representation is projected into a shared embedding space and used to condition an audio generation decoder. The interesting engineering is in the middle — the projection — because that is where the model either learns a real correspondence between visual and musical structure or collapses into producing the same pleasant loop for every input.

Alongside the model, the team is assembling an aligned image–music dataset and defining how this task should even be scored. Objective audio metrics say very little about whether a piece suits a picture, so evaluation combines retrieval-style measures with structured human listening comparisons.

Research questions

  1. Which visual attributes — palette and colour temperature, composition, density, subject matter — carry information the decoder can actually use?
  2. Do those attributes map onto musical ones in a stable way: tempo, mode, instrumentation, dynamic range?
  3. How much of the correspondence is genuinely learned, and how much is the model falling back on a prior over "generic pleasant music"?
  4. What evaluation protocol separates a musically competent output from an appropriate one?
Status
In progress
Started
10 August 2026 — ongoing
Modalities
Image → Audio
Areas
Multimodal AI · Generative Models
Researchers
Gurpreet Singh · Lamia Qamar
Script Panels
Fig. 2 — Segment narrative, plan layout, render panels
Project [10] Multimodal generation
In progress

Text-to-Comic Generation

Generating a single illustration from a sentence is close to solved. Generating twenty illustrations that tell one coherent story is not. This project takes written narrative as input and produces sequential comic pages — and the hard part is everything that has to stay true from one panel to the next.

The pipeline runs in stages. A script is segmented into narrative beats, one per panel; a layout planner decides how those beats sit on a page and in what reading order; an image generator renders each panel; and dialogue is extracted and placed into balloons positioned so the eye reaches them in the right sequence. Each stage is a research problem on its own, but the failures compound — a good panel in the wrong place still breaks the page.

The dominant open problem is consistency. A character must remain recognisably the same person across panels drawn independently, in changing poses, framing and lighting. The team is working on conditioning strategies that carry identity through the sequence, and on an evaluation that scores narrative faithfulness and character consistency separately from raw image quality.

The work builds directly on the lab's earlier storytelling and cross-modal generation projects, which handled narrative structure and modality translation independently.

Research questions

  1. How should continuous prose be segmented into discrete panels — by event, by dialogue turn, or by a learned boundary?
  2. What conditioning keeps a character identifiable across independently generated panels?
  3. Can panel layout and reading order be planned by the model rather than templated?
  4. How do you measure whether a generated page tells the story it was given?
Status
In progress
Started
Ongoing
Modalities
Text → Image sequence
Areas
Generative Models · Multimodal AI
Researchers
Trina Banerjee · Mukhthikka · Gurpreet Singh
Sequence Imaging Target edit
Fig. 3 — Sequence + imaging → targeted edit
Project [11] Multimodal · Biomedical
In progress

Multimodal AI in Gene Editing

Gene editing generates evidence in several forms at once — nucleotide sequence, microscopy and assay imaging, and a large published literature describing what has already been tried. Researchers routinely hold all three in mind together. Most computational tools handle one at a time.

This project applies the lab's cross-modal methods to that gap: encoding sequence, image and text evidence into a shared representation so a model can reason across them the way a biologist does. It is the same architectural question as our image-to-music and text-to-comic work — how do you build a space where two very different signals become comparable — pointed at a domain where being right matters a great deal more.

The near-term aim is modest and concrete: support target selection and outcome prediction, and be explicit about the confidence attached to each. The project sits squarely in the lab's AI-for-social-good area, alongside the earlier drug–drug interaction work.

Research questions

  1. Which representation lets sequence, image and literature evidence be compared without one modality dominating the others?
  2. Does adding imaging and literature actually improve target selection over sequence-only baselines, or only appear to?
  3. How should a model in this domain express uncertainty so that a biologist can act on it?
  4. What does responsible evaluation look like when the downstream application is biological rather than aesthetic?
Status
In progress
Started
Ongoing
Modalities
Sequence + Image + Text
Areas
Multimodal AI · AI for Social Good
Researchers
Sudipta Patil · Gurpreet Singh
Partners Shared work
Fig. 4 — Distributed sites, one research programme
Project [12] Cross-disciplinary
In progress

International Research Collaboration

Eight of the lab's members already work from outside India, and this project makes that distribution the subject rather than a logistical fact. It pairs the lab's AI methods with research in international business, ESG and media — the disciplines its members actually came from — and runs the work across time zones from the outset.

The substantive question is what AI changes about cross-border business and communication: how organisations report and are held to account, how markets and audiences behave across cultures, and where automated analysis genuinely helps rather than flattening the differences that matter. The methodological question is quieter but just as real — what a distributed research team needs in order to produce something coherent.

Scope and outputs are being defined with the collaborating partners; this entry will be expanded as they are fixed.

Research questions

  1. What can AI-driven analysis tell us about international business and ESG reporting that conventional methods miss?
  2. How do media narratives and consumer behaviour differ across markets, and can a model capture that without erasing it?
  3. Which parts of a cross-border research programme benefit from automation, and which depend on a person who knows the context?
  4. What working practice lets a distributed team sustain one line of enquiry?
Status
In progress
Started
Ongoing
Focus
International business · ESG · Media
Areas
AI Ethics & Bias · AI for Social Good
Researchers
Ritika Kanojia · Alimpia Roy · Gurpreet Singh

02 — Archive

Completed and closed work.

Sequence Segment boundaries
[08]

Neural Segmentation for Language Learning Analytics

Students developed and analysed neural segmentation models to identify and interpret linguistic patterns, in service of improving personalised language learning analytics.

2025.10.25 — 2026.01.10 Completed
Text Audio
[07]

Text-to-Audio (T2A) Generation

A cross-modal project generating audio from textual input across a range of deep learning techniques including CNNs, aimed at improving semantic understanding and audio synthesis in text-driven systems.

2025 Completed
Narrative branches
[06]

AI-Assisted Storytelling Generation

Integrating narrative structure, creativity and multimodal learning into automated storytelling systems. Published in the International Journal of Engineering Development and Research.

2025 Published
Vision Sound
[05]

Visual-to-Audio (V2A) Generation

Generating audio from visual input using computer vision, CNNs, RNNs and Vision Transformers — models that interpret a visual stimulus and produce a corresponding audio representation.

2025 Completed
Representation Discontinued
[04]

Analyzing Gender Bias in Media Using AI

A cross-cultural study of representation, language and imagery, combining NLP, computer vision and ethical frameworks. The project was discontinued before completion.

2025.07.21 — 2025.12.25 Cancelled
Growth analytics
[03]

Business Growth and Management Research Team

A team formed to explore the intersection of business growth, management and AI. Simran Singh, Nahid Rehmani, Pushpa Kumari and Vedha Varshini applied AI-driven analytics to business strategy and innovation.

2023 — 2024 Completed
Port surface :22
[02]

Cybersecurity and AI: Vulnerabilities in Modern Infrastructure

Gurpreet Singh and Prof. Saurabh Singh analysed the impact of open SSH ports on CentOS 9 virtual machines hosted on Mac ARM hardware (CVE-2023-25136), with recommendations for securing virtualised environments.

Spring 2024 Completed
Prompt exchange
[01]

ChatGPT and Cybersecurity: Implications and Repercussions

With Prof. Preethi Ananthachari, Gurpreet Singh examined the effect of large language models on cybersecurity — the opportunities and the risks generative AI introduces to digital security and privacy. Published in a leading journal.

2022 Published

Work with us

All four projects are taking members.

New members join one of the four running projects on arrival, with a mentor and a defined scope. If any of them is the kind of problem you want to spend months on, write to us.