Comparing Chameleon with GPT-4V and Gemini
The Chameleon model demonstrates new capabilities in mixed-modal understanding and generation. This section outlines human evaluations on large multi-modal language models' responses to diverse prompts commonly encountered by users. We describe our…
The Chameleon model demonstrates new capabilities in mixed-modal understanding and generation. This section outlines human evaluations on large multi-modal language models' responses to diverse prompts commonly encountered by users. We describe our methodology for prompt collection, evaluation baselines, and results. A safety study is also included.
To gather diverse prompts, we collaborated with a third-party crowdsourcing vendor. Annotators were asked to consider scenarios requiring multi-modal output, such as designing a kitchen layout. The prompts vary between text-only and mixed-modal (text and images). After collection, prompts were reviewed by three independent annotators for clarity and mixed-modal expectations. The final set consists of 1,048 prompts, with 42.1% being mixed-modal.
We compared Chameleon 34B with OpenAI GPT-4V and Google Gemini Pro. Responses from these models were text-only, so we enhanced them by generating image captions with DALL-E 3, creating GPT-4V+ and Gemini+ baselines. Evaluations were conducted using two approaches: absolute and relative.
The Chameleon model demonstrates new capabilities in mixed-modal understanding and generation.
Absolute Evaluation In absolute evaluations, each model's output was assessed by three annotators for task fulfillment. Chameleon's responses were deemed to fully fulfill tasks in 55.2% of cases, compared to 37.6% for Gemini+ and 44.7% for GPT-4V+. The task fulfillment varied across categories, with Chameleon excelling in Brainstorming and Comparison , while needing improvement in Identification and Reasoning .
Relative Evaluation Relative evaluations involved direct comparison of Chameleon against baseline models. Chameleon's responses were preferred 41.5% of the time over Gemini+, with a 35.8% win rate over GPT-4V+. Without image enhancements, Chameleon's superiority was more pronounced, with 53.5% over Gemini and 46.0% over GPT-4V.
This paper is available on arXiv under the CC BY 4.0 DEED license.
Based on reporting by hackernoon.com.
