Visual Question Answering with Frozen Large Language Models
Talking with LLMs about images, without training LLMs on images.
Who is this useful for? Data scientists interested in computer vision, natural language processing, and multimodal modeling.
How advanced is this post? Intermediate. You might struggle if you don’t have some experience in both computer vision and natural language processing.
Prerequisites: High level familiarity with transformers, embeddings, and encoder-decoders. All of these topics are covered in the following article:
A Brief Chronology of Visual Language Modeling
Visual language modeling really started up in 2016 with the paper VQA: Visual Question Answering, which formally posed the following class of problem:
Given an image and a natural language question about the image, the task is to provide an accurate natural language answer – VQA: Visual Question Answering
In 2016, when VQA was popularized, a typical approach looked something like this:
As vision and language models became more powerful, Visual Question Answering gave way to Visual Language Modeling (VLM), which can generally be considered as an expansion on visual question answering. Instead of simple questions like "is there a car in this image", modern Visual Language Models allow you to ask what type of car is in an image, then ask about how the car drives, the most popular movie that car was in, etc.
This shift from VQA to VLM was largely the result of incorporating large language models into visual systems, providing complex reasoning abilities and encyclopedic knowledge out of the box.
The difficulty of visual language modeling is, and always has been, multi-modality. You have to be good at images, natural language, and you have to be good at getting them to play nicely together.
The Q-Former in a nutshell
If you wanted to make a VQA system from scratch in a weekend, you might consider the following approach:
- Pass the image you want to talk about through a caption generator
- Combine the question asked by the user and the generated caption into a prompt for an LLM using some template
- Pass that prompt to the LLM, which would return the final output
The Q-former is used as a querying transformer (hence the name) which can transform a users query based on the image. The idea is to be able to extract the correct information from the image, based on the users prompt, and provide it to the LLM.
The BLIP-2 Architecture
Before we really dive into it, let’s get a high level understanding.
The BLIP-2 Architecture, which the Q-Former exists within, has the following components:
- An Image Encoder: A pretrained model which embeds images into an abstract representation which makes tasks like image classification easier. A popular example of this is CLIP.
- A Text Encoder: A pretrained model which embeds text into an abstract representation. A popular example of this is Word2Vect.
- An LLM: A large language model trained to perform general language tasks.
- The Q-Former: A transformer model which combines the embedded image and the embedded prompt into a format compatible for the LLM.
How the Q-Former Is Trained
The Training of the Q-Former can be divided into two phases: Bootstrapping and Generative Learning Pre-Training. The bootstrapping phase can be further divided into three sub phases:
- Image-Text Contrastive Learning: The model learns how to group image-caption pairs which belong together, and separate image-caption pairs which don’t belong together.
- Image-Grounded Text Generation: Divide the caption into two sections, the hidden and not hidden part, and attempt to guess the hidden part based on both the not hidden part and the image.
- Image-Text Matching: Pass the output of the Q-Former into a sacrificial dense network, which converts the output into a binary classification.
What We Get Out of Bootstrapping
Through the process of optimizing the Q-Former for these various tasks, the Q-Former is encouraged to build strong representations of both image and text, and a strong system for inter-relating the two.
VQA using Q-Formers from Hugging Face
In a future post I’ll be coding up and training a Q-Former from scratch, but for now let’s experiment with Q-Formers using a pre-built solution.
"""Downloading the BLIP-2 Architecture
loading as an 8 bit integer to save on GPU memory. This may have some impact on performance.
"""
from transformers import AutoProcessor, Blip2ForConditionalGeneration
import torch
processor = AutoProcessor.from_pretrained("Salesforce/blip2-opt-2.7b")
model = Blip2ForConditionalGeneration.from_pretrained("Salesforce/blip2-opt-2.7b", device_map="auto", load_in_8bit=True) # load in int8
Conclusion
In this post we went over the history of multimodal image and language modeling; from its humble beginnings in visual question answering to its modern stage of using large language models and image encoders.