Skip to content
All research

Research

ATAC-Qwen-VL

An accuracy-first image-captioning model for synthetic data, trained on consumer hardware.

Introduction

In this technical blog post, I am introducing my new model, AdapterTuneAdvancedCaptioning-QwenVL. It is designed specifically for creating synthetic image descriptions for large-scale datasets. I initially began building the model as a challenge to myself after LAION released their original GPT-V dataset [1]. My goal was to create a model with image captioning capabilities comparable to the original GPT4-V [2], using only personal resources. This means all my experiments and training were conducted on a personal GPU workstation. I achieved this by leveraging a self-QA training loop. This loop broke down the complex task of correcting incorrect image descriptions into many simple Visual QA tasks. I combined this training loop with the strong generalization ability of adapter-based fine-tuning strategies [3].

This combination allowed me to achieve the same level of detail and accuracy in image descriptions as GPT4-V. There have been several open-source and closed-source multi-modal releases, such as GPT4-V, LLava [4], and Phi-3-vision [5], all designed for human-centric interactions. In contrast, my model is not an instruction-based model. It is purely an image captioning tool designed for the automatic creation of large-scale, high-quality visual image datasets. These datasets can be used for pre-training image generation models like DALLE-3 [6]. This focus on synthetic data generation is also reflected in the writing style of the image descriptions. While other models produce image descriptions that are optimized to sound explicitly nice to humans, my model eliminates all "human-centric fluff". Instead, it outputs image descriptions in a shorthand-like format, similar to a Stable-Diffusion prompt [7].

Eval Image 1 Eval Image 2 Eval Image 4 Eval Image 5
GPT4-V
(gpt-4-1106-vision-preview)
This is an image of a female character who appears to be an adventurer or explorer, depicted in a highly detailed and realistic style that suggests she may be from a video game or a piece of concept art. The character is shown in a dynamic pose, with her left arm extended forward, holding a knife with a wrapped handle. Her expression is intense and focused, with her mouth slightly open and her brows furrowed, suggesting she is either in the midst of action or ready for an imminent threat.

She has a muscular build, with visible definition in her arms and shoulders, indicating strength and physical fitness. Her skin is tanned and has multiple scratches and smudges of dirt, implying she has been through rough conditions. Her clothing consists of a beige tank top, which is damp and clings to her torso, revealing her athletic physique. She wears brown cargo shorts with a belt and several pouches attached, which are practical for carrying equipment. Her outfit is completed with fingerless gloves and a pair of sturdy boots.

Around her neck, she wears a simple necklace with a pendant, and her dark hair is messy and damp, partially sticking to her face and neck, suggesting humidity or recent exertion. The character carries a large, beige backpack with visible wear and tear, hinting at extensive use.

The background is a lush, green jungle with dense foliage and a soft focus, creating a sense of depth and the impression that the character is deep within a tropical environment. The lighting in the scene is diffuse, with no harsh shadows, which could indicate an overcast sky or the canopy diffusing the sunlight.

Overall, the image conveys a sense of adventure, danger, and the resilience of the character within a wild and untamed natural setting.
This is a digitally created image featuring a woman in a dynamic archery pose. The woman is positioned in profile, facing to the left of the image, with her body slightly turned towards the viewer. She is drawing a modern compound bow with her right hand, while her left hand is holding the bow's grip. The bow is black with intricate designs and is equipped with a pulley system, a sight, and a stabilizer, indicating that it is designed for precision shooting.

The woman appears to be athletic, with toned arms and a strong posture. She is wearing a fitted, short-sleeved pink top that accentuates her physique, and olive green shorts that end mid-thigh. A quiver with several arrows with green fletching is strapped to her waist, and she has a single arrow nocked and ready to shoot. Her blonde hair is pulled back into a high ponytail with a colorful hair tie, and she is wearing stud earrings.

The background depicts a tropical beach scene with clear blue skies, fluffy white clouds, and a calm sea. The waves are gently crashing onto the shore, creating white foam. The sunlight appears to be casting a soft glow on the woman and the surrounding environment, suggesting it might be early morning or late afternoon.

The image is highly detailed, with attention given to the textures of the woman's clothing, the reflective surfaces of the bow's components, and the individual strands of her hair. The overall composition conveys a sense of focus, strength, and the serene beauty of the natural environment.
The image depicts a serene alpine landscape. In the foreground, there is a clear, calm mountain lake reflecting the blue sky above. The lake is surrounded by patches of green grass and rugged rocks, suggesting a high altitude environment. On the left side of the lake, there is a small green tent pitched on a flat area, indicating a temporary campsite. A wooden cross, possibly a trail marker or a memorial, stands on the right side of the lake, near the water's edge.

In the background, dominating the scene, are majestic snow-capped mountains. The peaks are rugged and sharp, with glaciers and snowfields clinging to their slopes. The highest peaks are shrouded in snow and ice, while the lower parts of the mountains show exposed rock and sparse alpine vegetation. The sky is mostly clear with a few wispy clouds, suggesting a beautiful, sunny day. The overall impression is one of tranquility and natural beauty, typical of a high mountain environment, possibly in the European Alps or a similar range.
The image depicts two individuals who appear to be in a warm, comfortable setting, likely at home. They are sitting at a wooden table with a laptop open in front of them, suggesting they might be working, studying, or browsing together.

On the left, there is a person with short, reddish-brown hair, wearing a white t-shirt and a smile, looking down at the laptop screen. Their posture is relaxed, and they seem to be engaged in what's on the screen. This person's ears are adorned with small earrings, and they exude a casual, content vibe.

On the right, there is another person with long, greenish-black braided hair, wearing a cozy, oversized pink sweater. They are holding a glass of what appears to be milk and leaning in towards the person on the left with a joyful expression, as if sharing a light-hearted moment or a laugh. Their nails are painted green, which adds a pop of color that matches their vibrant personality.

The table is adorned with a few items that suggest a comfortable, domestic setting: a grey mug, a clear jar with a metal clasp containing what looks like cookies, and a small plate with several oranges, adding a fresh, healthy touch to the scene.

The room is filled with natural light, which streams in from a window out of frame, casting a soft glow on the scene and creating a warm, inviting atmosphere. The overall mood of the image is one of companionship, comfort, and shared enjoyment in a relaxed home environment.
GPT4-V No fluff
(gpt-4-1106-vision-preview)
This is an image of a female character, depicted in a highly detailed and realistic style that. The character is shown in a dynamic pose, with her left arm extended forward, holding a knife with a wrapped handle. Her expression is intense and focused, with her mouth slightly open and her brows furrowed. She has a muscular build, with visible definition in her arms and shoulders. Her skin is tanned and has multiple scratches and smudges of dirt. Her clothing consists of a beige tank top, which is damp and clings to her torso, revealing her athletic physique. She wears brown cargo shorts with a belt and several pouches attached. Her outfit is completed with fingerless gloves and a pair of sturdy boots.

Around her neck, she wears a simple necklace with a pendant, and her dark, messy, and damp hair is partially sticking to her face and neck. The character carries a large, beige backpack with visible wear and tear. The background is a lush, green jungle with dense foliage and a diffuse focus. The lighting in the scene is diffuse, with no harsh shadows.
This is a digitally created image featuring a woman in a dynamic archery pose. The woman is positioned in profile, facing to the left of the image, with her body slightly turned towards the viewer. She is drawing a modern compound bow with her right hand, while her left hand is holding the bow's grip. The bow is black with intricate designs and is equipped with a pulley system, a sight, and a stabilizer.

The woman appears to be athletic, with toned arms and a strong posture. She is wearing a fitted, short-sleeved pink top that accentuates her physique, and olive green shorts that end mid-thigh. A quiver with several arrows with green fletching is strapped to her waist, and she has a single arrow nocked and ready to shoot. Her blonde hair is pulled back into a high ponytail with a colorful hair tie, and she is wearing stud earrings.

The background depicts a tropical beach scene with clear blue skies, fluffy white clouds, and a calm sea. The waves are gently crashing onto the shore, creating white foam. The sunlight appears to be casting a soft glow on the woman and the surrounding environment.

The image is highly detailed, with attention given to the textures of the woman's clothing, the reflective surfaces of the bow's components, and the individual strands of her hair.
The image depicts a serene alpine landscape. In the foreground, there is a clear, calm mountain lake reflecting the blue sky above. The lake is surrounded by patches of green grass and rugged rocks. On the left side of the lake, there is a small green tent pitched on a flat area, indicating a temporary campsite. A wooden cross, possibly a trail marker or a memorial, stands on the right side of the lake, near the water's edge.

In the background, dominating the scene, are majestic snow-capped mountains. The peaks are rugged and sharp, with glaciers and snowfields clinging to their slopes. The highest peaks are shrouded in snow and ice, while the lower parts of the mountains show exposed rock and sparse alpine vegetation. The sky is mostly clear with a few wispy clouds. The overall impression is one of tranquility and natural beauty, typical of a high mountain environment.
The image depicts two individuals who appear to be in a warm, comfortable setting, likely at home. They are sitting at a wooden table with a laptop open in front of them.

On the left, there is a person with short, reddish-brown hair, wearing a white t-shirt and a smile, looking down at the laptop screen. Their posture is relaxed, and they seem to be engaged in what's on the screen. This person's ears are adorned with small earrings.

On the right, there is another person with long, greenish-black braided hair, wearing a cozy, oversized pink sweater. They are holding a glass of what appears to be milk and leaning in towards the person on the left with a joyful expression, as if sharing a light-hearted moment or a laugh. Their nails are painted green.

The table is adorned with a few items that suggest a comfortable, domestic setting: a grey mug, a clear jar with a metal clasp containing cookies, and a small plate with oranges.

The room is filled with natural light, which streams in from a window out of frame, casting a soft glow on the scene.
AdapterTuneAdvancedCaptioning-QwenVL
(my Model)
A woman stands in the foreground of the image, her body positioned slightly to the left, with her right arm raised and her left arm down, Her right hand is gripping a knife with a dark, weathered blade, while her left hand is not visible, Her expression is intense, with a determined gaze and a serious, focused expression, Her hair is tousled and disheveled, with strands falling over her shoulders and to the sides, She wears a light gray tank top, with a darker gray or brown belt cinched around her waist, and a matching pair of shorts, Her right hand is adorned with a dark-colored glove, while her left hand is not visible, Her skin is covered in a layer of dirt and grime, with a visible scar on her left arm, The background is blurred, with a dense forest in the distance, suggesting a natural environment, The lighting is dramatic, with a strong contrast between the woman's dark, intense expression and the soft, diffused light that highlights her features, The image is realistic, with a high level of detail and a realistic style that emphasizes the woman's physicality and the environment around her A woman is positioned in the foreground, her body facing the ocean, with her right arm extended to the side, holding a bow and aiming it towards the water, Her left arm is bent at the elbow, holding the bowstring, and her right hand is gripping the bow, The bow is positioned in front of her, with the string pointing towards the water, She is wearing a pink sleeveless top, revealing her upper body, and dark green shorts, with a black belt, Her hair is styled in a ponytail with a colorful hairband, and she is wearing black wristbands, Her right hand is holding the bow, and her left hand is holding the bowstring, The background is a beach scene, with the ocean visible in the distance, and the sky with clouds, The waves are crashing against the beach, creating a misty effect, The woman's expression is focused and determined, with a serious look on her face, The realistic style of the image is enhanced by the high detail and the realistic lighting, which highlights the woman's features and the bow's details, The image is set against a realistic beach backdrop, with the ocean and sky in the background, The woman's attire is casual and athletic, with a focus on functionality and comfort, The realistic style of the image is evident in the high detail and the realistic lighting, which enhances the image's realism and the woman's athletic appearance A serene lake is nestled in the foreground, its calm surface reflecting the clear blue sky above, The mountain range, with its snow-capped peaks, is situated in the background, its rugged terrain and jagged peaks contrasting with the smooth, flat expanse of the lake, The mountain range is partially obscured by the lake, with its highest peaks visible above the water's surface, The mountain range's lower slopes are covered in green vegetation, rocks, boulders, The clear blue sky is reflected in the lake's still waters, creating a mirror-like effect, The landscape is devoid of any visible watermarks or text, emphasizing the natural beauty of the scene, The photograph's focus is sharp, with the mountain range and lake in clear focus, The natural lighting is bright, with the sun casting a warm glow on the landscape, The image's realistic style is evident in the sharp focus and the natural colors of the landscape, which are enhanced by the clear blue sky and the natural lighting, The photograph's composition is balanced, with the mountain range and lake taking up the majority of the frame, and the clear blue sky providing a backdrop, The image's serene atmosphere is created by the natural beauty of the landscape and the absence of any visible text or watermarks, which allows the viewer to fully appreciate the natural beauty of the scene In the image, two women are seated at a wooden table, the one on the left is wearing a white t-shirt and has short hair with a shaved part, while the one on the right is wearing a pink sweater with long green dreadlocks, The woman on the right is holding a glass of milk, which is positioned in front of her, and she is smiling broadly, The woman on the left is also smiling, with her hand resting on the other woman's shoulder, The table is adorned with a glass jar containing cookies, a jar with a lid containing a snack, and a clear glass mug, The window in the background provides natural light, casting shadows on the wall and table, The women's nails are painted in bright colors, with the woman on the right's nails adorned with green and purple hues, The image is in a candid and intimate style, with the women's expressions and body language conveying a sense of joy and companionship, The soft, diffused lighting suggests a bright and well-lit environment, with no harsh shadows or reflections, The focus is on the women, with the background and table items in soft focus, The image is free of text, brands, or watermarks, and the women's attire is casual and comfortable, with no specific clothing items or accessories mentioned
Figure 1: Comparison table for images between the original GPT-4V and AdapterTuneAdvancedCaptioning-QwenVL. Underlined text represents human-centric fluff from GPT4-v that does not add value to the image description. Bold text indicates statements about the image that the other model fails to capture. Red text indicates statements that are incorrect.

Failure Cases of Image Captioning Models

There are two main reasons for an Image captioning model to produce incomplete Image Descriptions. The first one is simply a writing style problem, where the

model mimics the style of pre-training data scraped from the internet, which is composed largely of shorter descriptions which only cover sub-aspects of the image.

Luckily, this is a fairly easy problem to fix, because the model still has learned all of the important underlying Image Text representations. It is just not outputting

them. To fix that, we simply have to fine-tune the model on a small dataset of more complete Image Descriptions, to bring the model to adapt a more detailed image

Description style.

Hallucinations

The main reason for hallucinations is a too low a resolution during the pre-training stages. The pre-training for both the Open-CLIP-bigG Image encoder [8] and the qwen-VL pre-training [9] were performed with an image resolution of 224x224 pixels. (qwen-VL had a later second pre-training stage where the image resolution was doubled to 448x448 pixels). Such a low resolution might be enough to capture the general composition of the image, but is inadequate to capture smaller details of the image.

Original image at 1024 by 1024 pixels

Origenal Image in 1024x1024 Pixels

Image scaled down to 64 by 64 pixels and enlarged

Image scaled down to 64x64 Pixels

Figure 2: Example of how information in the image is lost when downscaling the image.

For example, in the original image, the necklace is clearly visible, while in the scaled-down version, it is no longer discernible. A problem arises when the training

image descriptions still contain the information that the woman is wearing a necklace. This, combined with the fact that gradient descent is a local optimization

strategy, meaning it simply tries to find statistical correlations between the input data and the training target. In the best case, the statistical correlations make sense,

and the model learns what a necklace is and how it looks. In the worst case, gradient descent still tries to find some statistical correlations between the input and

target output, but they are completely nonsensical, like determining if a person is wearing a necklace by looking at the person's hairstyle or the lighting in the

background. When the model has learned these nonsensical correlations and is used to describe a new image, it might still predict that the person is wearing a

necklace, despite the person not actually wearing one. Fortunately, these correlations are usually more weakly connected within the model. Because of this, we can

easily remove them by using a negative training feedback strategy, such as RLHF [10] or DPO [11].

Diagram of the ATAC-Qwen-VL model architecture
Figure 3: Diagram of model architecture, trainable components and information flow.

Image Embedding strategy
I am introducing a Multi View Multi Crop Image Chunking Embedding Strategy (e.i. MVMCICE Strategy) which chunks an image into multiple crops at multiple

resolutions. Each image is embedded 6 times, at different resolutions.

Diagram of the multi-view multi-crop image chunking strategy
Figure 4: Diagram of the Multi View Multi Crop Image Chunking Embedding strategy. A single image is split into six chunk images at different resolutions.

These changes originate from a desire to increase the overall image resolution, together with the inability to perform an expansive image size adaptation of the Image Encoder. Naively increasing the image resolution would have led to multiple problems. The first one is that Qwen-LV [9] uses a fixed learned Query Attention mechanism to compress the image embeddings into a constant 256 tokens, regardless of the image input resolution. This severely limits the amount of information that the Image Encoder is able to pass to the underlying LLM. Had I naively increased the image resolution, I would have also wanted to increase the number of Embedding Tokens the Image Encoder is allowed to use. Both the change in image input size and re-training the learned Downsampling Queries, would have necessitated an extensive image size adaptation in order to ensure that the long information tail of the model is not severely degraded, which would have required orders of magnitude more compute than initially planned for this project. On the other hand, using the MVMCICE strategy has several upsides. It increases the overall image resolution from 448x448 pixels to 896x896 pixels. This, along with DPO training [11], helps alleviate the hallucination problem because the model becomes more capable of recognizing smaller objects in the image. The second big advantage is that I get embeddings per image covering a sub-chunk of the image. This is helpful because he embeddings also tend to ignore or fail to embed smaller details. But when the image is also embedded into sub-chunks, then there are fewer objects in the chunk and so even less important details get embedded. This helps to negate one of the primary reasons image captioning models produce incomplete image descriptions.

There are three main decisions/components to the model architecture, the first one is the Image Embedding strategy, the second one is the decision on which parameters to unfreeze, and the third is the selection of the pre-trained base model on which I am building on top of.

What is the Long Tail and Catastrophic Forgetting

But before we can take a look at my decisions, it makes sense to do a quick detour to some underlying training dynamics.
One important concept in deep learning models is the "long tail." If there are enough instances of a given concept in a training set, LLMs can learn robust and parameter-efficient representations of this concept. (In this context, parameter-efficient representations mean that the model can effectively learn common concepts using fewer parameters, while more parameters are required to learn and represent rare concepts.) On the other hand, when there are just a few instances of a given concept, the model needs lots of parameters to learn it. This is called the long tail [12]. One significant reason why the long tail is important is that it helps the model to generalize better. When a model is exposed to a wide variety of rare concepts, it becomes more adept at understanding diverse constructs. This improves the model's ability to robustly handle a broad range of inputs.
Catastrophic forgetting occurs when a model is trained sequentially on multiple tasks or datasets. As the model adapts to new tasks, it tends to "forget" previously learned information, concepts, and patterns. Not all concepts degrade at the same rate during catastrophic forgetting. The first information to go is often the rare concepts contained in the model's long tail. This dramatically degrades the model's ability to generalize. This issue becomes critical when we have a small dataset and a distribution shift between training data and actual use cases. In such scenarios, the model must still generalize well to perform effectively during inference [13].

Introduction to Finetuning Strategies

Full Model

Full model fine-tuning involves updating all the parameters of a pre-trained model with new data. This process provides the most adaptability for learning new information. However, it comes with the cost of catastrophic forgetting, where the model loses performance on previously learned tasks due to comprehensive parameter updates. I would generally not recommend using full model fine-tuning outside of pre-training dataset scale where the model adapts to a completely new domain, with the expectation that it will lose some of its world knowledge from the previous pre-training. A notable example is the large-scale image pre-training phase that the Qwen-VL model underwent [9]. If the new data significantly differs from the old training objectives, the model may encounter problematic gradients early in the training process, leading to a significant degradation in performance. This issue can arise, for example, during the image encoder adaptation for multi-modal training. To overcome this problem, there are multiple approaches. Qwen-VL employs a multi-stage training setup. Initially, only the image encoder is updated while the base LLM remains frozen. This allows the image encoder to learn an embedding space compatible with the LLM's input space. In the second stage, all the parameters of both the image encoder and the LLM are updated. Meanwhile, MAGMA [14] incorporates adapter parameters, training them alongside the image encoder parameters without unfreezing the original LLM parameters or performing a second pre-training stage.

LoRA

Low-Rank Adaptation (LoRA) [15] introduces trainable low-rank matrices to each layer of a pre-trained model while keeping the original weights frozen. Typically, LoRAs are applied only to the query and value projections within the attention mechanism, significantly reducing the number of trained parameters in comparison to full model fine-tuning. The LoRA weights can be fully fused with the original parameters, adding no extra latency during inference. LoRA is a highly popular fine-tuning strategy due to its low memory footprint and Hugging Face's out-of-the-box support for most models. Another advantage of LoRAs is their small memory footprint, allowing for hundreds or even thousands of different LoRAs to be loaded onto the same GPU and hot-swapped during inference.
However, there are several downsides associated with LoRA fine-tuning. Due to the limited number of trainable parameters, LoRA training is less adaptable than full model training. Additionally, since LoRA modifies some original parameters, it can lead to catastrophic forgetting, although less severe than with full model fine-tuning. It also results in reduced performance when there is a distribution shift between the training data and the inference use case. Despite its popularity in the ML community, I would not recommend using LoRA as the default unless there is a requirement for the lowest possible inference latency or it is explicitly deployed in a setting where many LoRAs are simultaneously kept on a single GPU.

Adapters

Adapters [3] are small MLP blocks added within the Attention and MLP blocks of pre-trained Transformer models. During fine-tuning, only the adapter parameters are updated, while the original model parameters remain untouched. This approach ensures that the fine-tuning strategy does not suffer from catastrophic forgetting. As a result, adapter fine-tuned models tend to outperform full model fine-tuning and LoRAs when a distribution shift occurs between the training data and the actual use case. (A distribution shift can be subtle, such as training on general documents for a QA task and then applying the model to a more specific context like medical QA.) Adapters remain quite parameter-efficient compared to full model fine-tuning, though not as much as LoRAs. Since additional small MLP blocks are added to the base model, they can easily be disabled or swapped with other adapters during inference. Another downside of adapters is that the addition of small MLP blocks increases inference latency. Adapters can handle more data than LoRAs due to their higher parameter count, but not as much as full model fine-tuning. Since none of the original parameters are updated, adapter-based fine-tuning avoids catastrophic weight updates, making it a stable fine-tuning strategy. Overall, I recommend it as the default fine-tuning strategy. Unfortunately, adapters require changes to the model architecture. As such, they are not supported by default by Hugging Face or other popular inference libraries, making them less popular.

SoftPrompt

SoftPrompt tuning [16] involves learning a set of soft prompt vectors, which are prepended to the input sequence of a pre-trained model. SoftPrompt fine-tuning does not update the model's weights; instead, it optimizes the prompt vectors to guide the model in performing a new task. One benefit is that SoftPrompts have a hard time overfitting, even after hundreds of epochs, because they can only steer a language model in a fairly subtle way. This allows SoftPrompts to be trained on datasets with fewer than 10 samples. The downside of SoftPrompts is that the model must already have some ability to perform the given task, or it must be combined with other fine-tuning strategies.

Available Base Model

Over the years, many multimodal models have been made publicly available. Most of these models are based on pre-trained LLMs with an added image encoder and optional additional parameters, while the LLM remains frozen. The original MAGMA model [14] used CLIP [17] as its image encoder and added adapters to a frozen, pre-trained LLM. The publicly available MAGMA model was based on the 6B GPT-J model [18] and was trained on 7.6M image-text pairs. Flamingo [19] (not publicly available) and OpenFlamingo [20] build on this idea, using cross-attention layers instead of adapters and directly feeding the CLIP embeddings [17] into the LLM. Flamingo was trained on 2.3B image-text pairs based on a 70B model, while OpenFlamingo was trained on 180M image-text pairs and based on a 7B model.

Another popular multimodal model is BLIP-2 [21]. The largest version of BLIP-2 has 12.1B parameters and is based on a FlanT5xxl model and a ViT-g image encoder. It was trained on 129M image-text pairs. There are also two other noteworthy open-source models trained on large image-text pair datasets: CogVLM [22] and Qwen-VL [9].

CogVLM [22] is based on Vicuna1.5 7B [23], an instruct fine-tuned and RLHF version of Llama 2 7B [24], together with an EVA2-CLIP-E 4.4B image encoder [25]. CogVLM duplicates all the MLP and attention parameters and builds with it a fixed routed mixture of experts. The duplicated parameters are enabled on the token positions from the image encoder, while the original parameters are used for text tokens. The two parameter sets connect over the attention mechanism, routing information from the image tokens to the text tokens. During training, all the original parameters remain frozen, while only the duplicated parameters are trained. This prevents catastrophic forgetting of the base model while maintaining high adaptability for image training.

The final model I examined was Qwen-VL [9]. It is based on an early checkpoint of the Qwen 7B model [26] and uses the 2B Laion CLIP model [8] as its image encoder. The image encoder uses a set of 256 learned query embeddings, which downsample the image embedding token dimension to a fixed size of 256 token embeddings. These are then passed to the LLM. This makes the number of embedding tokens independent of the image resolution used. The model was trained in three stages:

  • The first stage involved adapting the image encoder embeddings, where the LLM remained frozen and only the image encoder was updated. This was likely done to prevent catastrophic gradients when unfreezing all weights during the second training stage.

  • The second stage introduced a large-scale bounding box segmentation task and OCR data-set, along with the standard image captioning task commonly used for training multimodal LLMs.

  • The third stage involved adapting the image size from 224x224 to 448x448 pixels.

Overall, the model was trained on 1.4B images.

My Choice

While writing the blog post, I realized I should have chosen the CogVLM model [22] instead of the Qwen-VL model [9]. I was aware of this at the start of the project, but I rejected it because I mistakenly believed it was a full model fine-tuning based on the 13B Stable Vicuna 1.0 [27], an early open-source instruction model derived from LLaMA1 [28]. In reality, it is based on Vicuna 7B 1.5 [23], which is built on LLaMA2 7B [24], with additional instruction and RLHF fine-tuning. This significantly improves performance over the 1.0 versions of the Stable Vicuna models. My second misconception was that I wrongly assumed it to be a full model fine-tuning of a 13B model, when in reality, it was a weight duplication of a 7B model, which keeps all original parameters frozen while adding new trainable parameters. This is highly advantageous as it prevents catastrophic forgetting while providing sufficient model capacity to absorb the 1.5 billion image-text pairs. Another reason I should have chosen the CogVLM model is that Qwen-VL uses learned query embeddings to downsample the number of image tokens to a fixed 256 tokens. Unfortunately, this creates an information bottleneck because the downsampling mechanism relies on the query learning the core concepts which are passed to the LLM. A clear advantage of Qwen-VL over CogVLM is that it has only 9B parameters, including the image encoder and LLM, compared to CogVLM's 17B parameters. This makes Qwen-VL easier to handle and trainable on 24GB GPUs. Another aspect I liked about Qwen-VL is its solid dataset mix, which includes not only image-text pairs but also bounding box annotations and extensive OCR data. This makes Qwen-VL quite proficient at reading text in images. I believe Qwen-VL was a fine base model, but I could have built a substantially better image captioning model with CogVLM.

For my fine-tuning, I primarily used adapters in addition to a small soft prompt. I mainly chose an adapter finetuning strategy because my datasets are relatively small. I had just over 12K samples for the STF and less than that for each preference alignment iteration. Due to the small number of samples, I needed to heavily rely on the base model's generalization capability, making it crucial to keep the base model's long tail intact. Adapters are deployed not only to the LLM but also to the transformer blocks of the ViT-based CLIP image encoder [8], allowing for its training as well. The Qwen-VL model was trained with a short prompt to distinguish between Chinese and English outputs and to indicate whether object bounding boxes are predicted in line with the image description. Therefore, I am using a short soft prompt initialized from the token embeddings of the English captioning task prompt without bounding boxes. I consider the addition of the soft prompt more of an implementation detail. I am confident that the adapters alone have enough capacity to learn the task as effectively as the combination of adapters and a short soft prompt.

Training Loop

There are three main steps involved in the training loop. The training of QA Adapters, the generation of a golden sample, and a Training step to update the Model.

Diagram of the ATAC-Qwen-VL training setup
Figure 5: Diagram of the training setup, the training process for the QA adapter, the creation of golden samples, and the subsequent model update.

Training of the QA Adapter

The first step is quite straightforward. I am training a QA model based on the current image Captioning Model in order to ask it questions about the Image. For that, I am initializing a new Adapter below the Captioning Adapter in only the LLM part of the model. No parameter in the CLIP Image Encoder is trained. This has the background that the QA dataset has a narrower distribution than the more general distribution from the Web Corpus used as Image source. Training parameters in the CLIP Image Encoder could lead to a collapse of the image embeddings towards the narrower QA dataset distribution.
The training set of the QA adapter is a small collection (around 2K images) of Image Question and Answer triplets, which are in the style that the question asks if some aspect/object is in the image and the answer then confirms or denies the statement.

Combination of Multiple Generations

In order to increase the information level of the image descriptions, I am generating multiple ones and then combining their content into a single one using a tool LLM. This helps because each description tends to include some different aspects from the image, together with some hallucinations. The resulting combined image description is significantly longer than any of the original image descriptions.

Decomposition into Simple Questions

The combined image description is now a lot longer than image descriptions from the model distribution but also contains some hallucinations. To get rid of the hallucination, I am using again a tool LLM to decompose the statements made in the image description into fine-grained questions about the image, which ask if the statement is true or not. Each image description then results on average in between 30 and 120 questions. Each question is given to the QA model which was trained in step one. The QA model is generally capable of telling if the statement is in the image and true or not, even while using the same underlying model. This has multiple aspects; first of all, it is a lot easier task to answer a direct question than to generate a full image description out of thin air. The second reason why the QA model is generally capable of correcting hallucinations is that most hallucinations are created by sampling with temperature and because of that are coming more from the weaker connections inside the model. The remaining error, which the QA model is not able to filter out, is a statistically acceptable error and will be removed by performing multiple training iterations. The long and de-hallucinated image description is considered a "Golden Sample" and will be used as the positive sample together with the in-model distribution image descriptions as negative samples.

Training

The Qwen-VL model [9] got supervised finetuned with a reformulated version of the LAION GPT4-V-dataset [1] containing around 12K Image Descriptions collected by GPT4-V. A LLM was used to reformulate the human-centric writing style to a more mechanical and shorthand style similar to the prompt style of Stable Diffusion Models [7]. For the supervised finetuning, all adapters in both the image encoder and the LLM were trained together with the embedding soft prompt. For all following preferences alignment iterations, the soft prompt stayed frozen and only the adapters in both the image encoder and LLM continued to be trained.
During the preferences alignment step, I encountered severe training instabilities in the form that the model tended to completely collapse and started to only repeat single nonsensical statements. The exact instability had a strong dependency on the used data shuffling seed, where the same hyperparameters with only different data seed would lead to dramatically different models. I also ablated over different dpo beta values, and over standard DPO [11], Hinge_DPO [29], IPO [30], and KTO [31]. From all of these combinations, only DPO with a batch smaller than 0.1 was able to produce some somewhat usable models 1 out of 4 times. All other Direct preferences alignment methods completely failed to produce a working model. I was able to stabilize the training by combining a DPO loss with a Cross entropy loss on the selected sample. This loss combination dramatically stabilized the training and made it also fairly seed to seed repeatable. My hypothesis on why DPO struggles is that the selected sample is quite far away from the normal model distribution, due to the multisample combination and QA de-hallucination steps.
Another common failure case after the supervised finetuning and the first preferences aligned iteration was that the model did not produce an EOS token and instead started to repeat the last made statement indefinitely. This failure case did disappear nearly completely after the second preferences alignment iteration.
I trained for three data generation and preferences alignment training iterations, where the third iteration resulted in a degradation of performance compared to the previous iteration. The published model is the resulting model from the second iteration. I hypothesize that the degradation of the third iteration has mainly the aspect that the image descriptions started to significantly push out of the context size of the pretrained model. Also, it is conceivable that the preferences data generation pipeline, or the used foundation model was pushed over its limit. I did not investigate the degradation of the third preferences aligned iteration or tried to fix it, because the second iteration was already a fairly powerful model, which reached my self-selected goal, of creating an image captioning model as good as GPT4-V.
All experiments, trainings, and preferences data generations were performed on a DIY Hexa 3090 workstation over the course of three months.

Eval Image 1 Eval Image 2
After pre-training Lara Croft Tomb Raider 2018 By Jacek Yerka Tomb Raider Tomb Raider 2018 Tomb Raider Movie Archery Girl
After finetuning digital illustration of a female character in a forest, holding a knife with a brown handle, wearing a sleeveless top, shorts, belt, gloves, backpack, muddy, wet, looking at the viewer, intense expression, short hair, brown eyes, light skin, jungle background with greenery, realistic style, detailed textures, lighting, no text digital illustration, woman aiming bow and arrow, sandy beach, ocean background, long blonde hair in ponytail, pink sleeveless top, olive green shorts, black quiver with arrows, black wristband, black wrist brace, bow with black compound, black riser, black string, black cable, black wheel, green arrow in mouth, focused expression, serene ocean waves, clear sky, no text
After Iteration 1 woman in outdoor setting, holding a knife with a dark blade, wearing a light gray tank top, short brown shorts, a brown belt with multiple pockets, a tan backpack with a shoulder strap, a necklace with a pendant, a black arm band, a bracelet on her right wrist, a black glove on her left hand, a dirty white bandage on her right knee, a black watch on her left wrist, a dark-colored hair tied back, looking directly at the camera with a serious expression, surrounded by greenery with a blurred background, realistic style with a focus on detail and texture, no text or brands visible, suitable for video games or promotional materials woman aiming with bow, light pink sleeveless top, blonde hair in a high ponytail, black hairband with colorful accents, black gloves on both hands, holding bow with right hand, left hand on bowstring, black quiver with arrows visible, black knee-length shorts, black ankle boots with heel, standing on a sandy beach, ocean visible in the background, clear blue sky with scattered clouds, realistic style with digital painting techniques, no text or brands visible, emphasis on athletic and artistic expression, female archer character, and the bow and arrow equipment
After Iteration 2
(Published model)
A woman stands in the foreground of the image, her body positioned slightly to the left, with her right arm raised and her left arm down, Her right hand is gripping a knife with a dark, weathered blade, while her left hand is not visible, Her expression is intense, with a determined gaze and a serious, focused expression, Her hair is tousled and disheveled, with strands falling over her shoulders and to the sides, She wears a light gray tank top, with a darker gray or brown belt cinched around her waist, and a matching pair of shorts, Her right hand is adorned with a dark-colored glove, while her left hand is not visible, Her skin is covered in a layer of dirt and grime, with a visible scar on her left arm, The background is blurred, with a dense forest in the distance, suggesting a natural environment, The lighting is dramatic, with a strong contrast between the woman's dark, intense expression and the soft, diffused light that highlights her features, The image is realistic, with a high level of detail and a realistic style that emphasizes the woman's physicality and the environment around her A woman is positioned in the foreground, her body facing the ocean, with her right arm extended to the side, holding a bow and aiming it towards the water, Her left arm is bent at the elbow, holding the bowstring, and her right hand is gripping the bow, The bow is positioned in front of her, with the string pointing towards the water, She is wearing a pink sleeveless top, revealing her upper body, and dark green shorts, with a black belt, Her hair is styled in a ponytail, with a colorful hairband, and she is wearing black wristbands, Her right hand is holding the bow, and her left hand is holding the bowstring, The background is a beach scene, with the ocean visible in the distance, and the sky with clouds, The waves are crashing against the beach, creating a misty effect, The woman's expression is focused and determined, with a serious look on her face, The realistic style of the image is enhanced by the high detail and the realistic lighting, which highlights the woman's features and the bow's details, The image is set against a realistic beach backdrop, with the ocean and sky in the background, The woman's attire is casual and athletic, with a focus on functionality and comfort, The realistic style of the image is evident in the high detail and the realistic lighting, which enhances the image's realism and the woman's athletic appearance
After Iteration 3
(Failed Iteration)
A woman with dark hair, styled in a natural, wavy manner, stands in the foreground of the image, her body positioned slightly to her right, with her right arm extended forward, holding a knife with a dark, possibly metallic blade, the knife is held in a grip position, with her right hand wrapped around the handle, her left hand is positioned in a similar grip position, with her fingers wrapped around the handle, her body is angled slightly to her right, with her right leg forward and her left leg slightly behind, her right arm is extended forward, and her left arm is bent at the elbow, with her hand resting on her hip, her expression is intense and focused, with a serious, determined look, her eyes are open, looking directly at the camera, her skin is covered in a layer of dirt, suggesting a rugged, outdoor environment, her attire includes a light gray tank top, a dark belt, and dark cargo shorts, with a dark belt and a dark, possibly leather or fabric, possibly canvas, shoulder strap, her right hand is wrapped around the handle of the knife, and her left hand is resting on her hip, her right arm is extended forward, and her left arm is bent at the elbow, her hair is tousled and appears to be wet, possibly from being in the water or from sweating, the background is a dense, green forest, with trees and foliage in the background, the lighting is bright and natural, with soft, diffused light, casting shadows on her face and body, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is in high resolution, with a clear focus on the woman and her surroundings, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is in high resolution, with a clear focus on the woman and her surroundings, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is in high resolution, with a clear focus on the woman and her surroundings, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is in high resolution, with a clear focus on the woman and her surroundings, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is in high resolution, with a clear focus on the woman and her surroundings, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's attire is functional and practical, with a focus on functionality and comfort, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the environment, the woman's expression is intense and focused, with a serious, determined look, the image is realistic, with a high level of detail and a realistic, detailed portrayal of the woman's attire and the woman's expression, the image is realistic, with a high level of detail and a realistic portrayal of the woman's attire and the woman's expression, the image is realistic, with a high level of detail and a realistic portrayal of the woman's attire and the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, with a high level of the woman's expression, the image is realistic, the image is realistic, the image is realistic, the woman's expression, the image is realistic, the image is realistic, the image is realistic, the image is realistic, the image is realistic, the image is realistic, the image is realistic, the image is realistic, the view, the image, the view, the view, the the view, the the view, the a view, the view, the view, the view, the view, the view, the view, the view, the the view, the view, the the view, the the view, the a view, the the the view, the the the the the view, the a, the a view, the a view, the view, the a, the the view, the a, the a, the the view, the a, the a, the a, the a, the a, the a, the a, the a, the a, the a, the a, the a, the a, the a A woman stands on a beach, her body positioned slightly to her right, with her right arm extended, holding a bow, the bow is positioned in front of her, with the string pointing towards her, and the arrow is attached to the bow, she is wearing a pink top, which is fitted and form-fitting, revealing her chest and shoulders, and a pair of dark green shorts, her hair is styled in a high ponytail, with a small, colorful hair tie, she is wearing black wrist guards, and her right hand is positioned on the bowstring, holding the bow, her left hand is holding the bow, with her fingers wrapped around the bowstring, the background is a beach scene, with the ocean visible in the distance, the sky is clear and blue, with a few clouds, the waves are visible, crashing against the beach, creating a misty effect, the woman's expression is focused and determined, with her eyes looking towards the target, the image is realistic, with a high level of detail and a realistic style, the lighting is bright and natural, with the sun casting shadows to the left, and the water is a light blue, with white foam, the image is free of text, watermarks, or other distractions, the woman's body is in sharp focus, with the bow and the background in soft focus, the image is clear and focused, with the woman's body and the bow in sharp focus, the woman's attire is athletic and functional, with the shorts and top designed for athletic activities, the image is realistic, with a high level of detail and a realistic style, the woman's body is in sharp focus, with the bow and the background in soft focus, the image is free of text, watermarks, or other distractions, the woman's expression is focused and determined, with her eyes looking towards the target, the image is realistic, with a high level of detail and a realistic style, the woman's body is in sharp focus, with the bow and the background in soft focus, the image is free of text, watermarks, or other distractions
Figure 6: The table showcases the changes in image descriptions generated by the model from the pre-trained stage through the third training iteration.

References

  1. LAION, “GPT-4V Dataset,” Hugging Face Datasets, 2023. [Online]. Available: huggingface.co/datasets/laion/gpt4v-dataset.
  2. OpenAI, “GPT-4V(ision) System Card,” September 25, 2023. [Online]. Available: openai.com/index/gpt-4v-system-card.
  3. N. Houlsby et al., “Parameter-Efficient Transfer Learning for NLP,” arXiv:1902.00751, 2019. [Online]. Available: arxiv.org/abs/1902.00751.
  4. H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual Instruction Tuning,” arXiv:2304.08485, 2023. [Online]. Available: arxiv.org/abs/2304.08485.
  5. Microsoft, “Phi-3-Vision-128K-Instruct Model Card,” May 21, 2024. [Online]. Available: huggingface.co/microsoft/Phi-3-vision-128k-instruct.
  6. J. Betker et al., “Improving Image Generation with Better Captions,” OpenAI, 2023. [Online]. Available: cdn.openai.com/papers/dall-e-3.pdf.
  7. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” arXiv:2112.10752, 2021. [Online]. Available: arxiv.org/abs/2112.10752.
  8. M. Cherti et al., “Reproducible Scaling Laws for Contrastive Language-Image Learning,” arXiv:2212.07143, 2022. [Online]. Available: arxiv.org/abs/2212.07143.
  9. J. Bai et al., “Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,” arXiv:2308.12966, 2023. [Online]. Available: arxiv.org/abs/2308.12966.
  10. L. Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback,” arXiv:2203.02155, 2022. [Online]. Available: arxiv.org/abs/2203.02155.
  11. R. Rafailov et al., “Direct Preference Optimization: Your Language Model Is Secretly a Reward Model,” arXiv:2305.18290, 2023. [Online]. Available: arxiv.org/abs/2305.18290.
  12. Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie, “Class-Balanced Loss Based on Effective Number of Samples,” arXiv:1901.05555, 2019. [Online]. Available: arxiv.org/abs/1901.05555.
  13. J. Kirkpatrick et al., “Overcoming Catastrophic Forgetting in Neural Networks,” arXiv:1612.00796, 2016. [Online]. Available: arxiv.org/abs/1612.00796.
  14. C. Eichenberg, S. Black, S. Weinbach, L. Parcalabescu, and A. Frank, “MAGMA—Multimodal Augmentation of Generative Models through Adapter-based Finetuning,” arXiv:2112.05253, 2021. [Online]. Available: arxiv.org/abs/2112.05253.
  15. E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv:2106.09685, 2021. [Online]. Available: arxiv.org/abs/2106.09685.
  16. B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-Efficient Prompt Tuning,” arXiv:2104.08691, 2021. [Online]. Available: arxiv.org/abs/2104.08691.
  17. A. Radford et al., “Learning Transferable Visual Models from Natural Language Supervision,” arXiv:2103.00020, 2021. [Online]. Available: arxiv.org/abs/2103.00020.
  18. B. Wang and A. Komatsuzaki, “GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model,” EleutherAI, 2021. [Online]. Available: huggingface.co/EleutherAI/gpt-j-6b.
  19. J.-B. Alayrac et al., “Flamingo: A Visual Language Model for Few-Shot Learning,” arXiv:2204.14198, 2022. [Online]. Available: arxiv.org/abs/2204.14198.
  20. A. Awadalla et al., “OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models,” arXiv:2308.01390, 2023. [Online]. Available: arxiv.org/abs/2308.01390.
  21. J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” arXiv:2301.12597, 2023. [Online]. Available: arxiv.org/abs/2301.12597.
  22. W. Wang et al., “CogVLM: Visual Expert for Pretrained Language Models,” arXiv:2311.03079, 2023. [Online]. Available: arxiv.org/abs/2311.03079.
  23. LMSYS Org., “Vicuna-7B v1.5 Model Card,” Hugging Face Models, 2023. [Online]. Available: huggingface.co/lmsys/vicuna-7b-v1.5.
  24. H. Touvron et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv:2307.09288, 2023. [Online]. Available: arxiv.org/abs/2307.09288.
  25. Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao, “EVA-CLIP: Improved Training Techniques for CLIP at Scale,” arXiv:2303.15389, 2023. [Online]. Available: arxiv.org/abs/2303.15389.
  26. J. Bai et al., “Qwen Technical Report,” arXiv:2309.16609, 2023. [Online]. Available: arxiv.org/abs/2309.16609.
  27. CarperAI, “StableVicuna-13B Model Card,” Hugging Face Models, 2023. [Online]. Available: huggingface.co/CarperAI/stable-vicuna-13b-delta.
  28. H. Touvron et al., “LLaMA: Open and Efficient Foundation Language Models,” arXiv:2302.13971, 2023. [Online]. Available: arxiv.org/abs/2302.13971.
  29. Hugging Face, “DPO Trainer: Loss Functions,” TRL Documentation. [Online]. Available: huggingface.co/docs/trl/dpo_trainer.
  30. M. G. Azar et al., “A General Theoretical Paradigm to Understand Learning from Human Preferences,” arXiv:2310.12036, 2023. [Online]. Available: arxiv.org/abs/2310.12036.
  31. K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela, “KTO: Model Alignment as Prospect Theoretic Optimization,” arXiv:2402.01306, 2024. [Online]. Available: arxiv.org/abs/2402.01306.

Enlarged article image