Bee Logo Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs

ICLR 2026  Β·  Computational Visual Media (journal version)

Journal version: Bee: A High-Quality Corpus and Full-Stack Suite for End-to-End Training of Fully Open MLLMs

Yi Zhang1,2, Bolin Ni2, Xin-Sheng Chen1, Heng-Rui Zhang1, Yongming Rao2, Houwen Peng2*, Qinglin Lu2, Han Hu2, Meng-Hao Guo1†, Shi-Min Hu1
1Tsinghua University, 2Tencent Hunyuan Team
*Project lead. †Corresponding author.

πŸ”₯ News

  • [2026.09] 🍯 Honey-Data-V2 is released: 44.3M samples comprising 100.3M QA pairs over 19.7M image groups, the scaled-up corpus of the journal version. See below.
  • [2026.09] πŸŽ‰ The journal version of Bee, Bee: A High-Quality Corpus and Full-Stack Suite for End-to-End Training of Fully Open MLLMs, has been accepted to Computational Visual Media.
  • [2026.03] πŸ’» Our main code is open-sourced on GitHub.
  • [2026.01] πŸŽ‰ Bee has been accepted to ICLR 2026.
  • [2025.12] πŸ”₯ All data and model weights across training stages are released.
  • [2025.11] πŸ“Š Honey-Data-15M and Honey-Data-1M are released.
  • [2025.10] 🐝 Bee-8B is released.

✨ What's New in the Journal Version

The conference version mainly focused on the SFT stage and treated RL in a relatively decoupled manner. The journal version considers SFT and RL in a unified end-to-end recipe:

  • Model-tailored RL data. Building on the original SFT corpus and its curation pipeline, we introduce a model-tailored RL data construction pipeline and a high-quality RL dataset, completing the full workflow from SFT to RL (below).
  • Honey-Data-V2. The SFT corpus itself is scaled up to 44.3M samples comprising 100.3M QA pairs: the original pool is re-curated under stricter structural constraints, re-annotated with an upgraded model stack, and extended with newly released community corpora (below).
  • A more comprehensive empirical study. Results previously distributed across the main paper and supplementary material are consolidated and extended with additional analyses, ablation studies, and stage-wise evaluations.

πŸ“„ Abstract

Fully open multimodal large language models (MLLMs) still trail proprietary counterparts, largely due to a gap in end-to-end training data quality. Existing open-source efforts often train individual stages rather than a unified end-to-end pipeline. At the same time, currently available datasets still suffer from pervasive noise, a shortage of complex reasoning data such as chains-of-thought (CoTs), and a lack of reinforcement learning (RL) signals well aligned with model capabilities.

To address these challenges, our work makes three primary contributions. First, we introduce Honey-Data-15M, a supervised fine-tuning (SFT) corpus of approximately 15 million QA pairs cleaned using multiple techniques and enhanced with a novel dual-level (short and long) CoT strategy. Alongside Honey-Data-15M, we also present a high-quality RL dataset specifically tailored to the model's capabilities. Second, we introduce HoneyPipe, a data curation pipeline, and its underlying framework DataStudio, providing the community with a transparent and adaptable methodology for full-pipeline data curation (pre-training, SFT, and RL) that moves beyond static dataset releases. Finally, to validate our dataset and pipeline, we have trained Bee-8B through a joint SFT and RL recipe.

Experiments show that Bee-8B establishes a new state-of-the-art for fully open MLLMs, achieving performance that is competitive with, and in some cases surpasses, recent semi-open models such as InternVL3.5-8B. A comprehensive ablation study further dissects the impact of our data curation process, revealing that each stage provides significant performance gains across a wide range of benchmarks. Our work delivers to the community a suite of foundational resources, including the Honey-Data-15M corpus, the full-stack suite comprising HoneyPipe and DataStudio, training recipes, an evaluation harness, and the model weights. This effort demonstrates that a principled focus on end-to-end data quality is a key pathway to developing highly competitive fully open MLLMs.

βš™οΈ The HoneyPipe Pipeline

To address the challenges of data noise and the reasoning gap in open-source datasets, we developed HoneyPipe, an automated and reproducible workflow built on our DataStudio framework. It systematically transforms a vast, raw data pool into a high-quality, dual-level Chain-of-Thought (CoT) dataset suitable for supervised fine-tuning (SFT).

HoneyPipe Pipeline Stages

Figure 1: The HoneyPipe data curation pipeline with five key stages: data aggregation, noise filtering, short CoT enrichment, long CoT enrichment, and fidelity verification.

The pipeline consists of several key stages:

  • Data Aggregation and Deduplication: We start by assembling ~24 million image-text pairs from diverse sources and perform rigorous deduplication to maximize data diversity and processing efficiency.
  • Noise and Irrelevance Filtering: This stage uses both rule-based and model-based operators to purge noisy data, removing samples with formatting issues, low-quality images, or image-instruction mismatches.
  • Short CoT Enrichment: For instructions requiring moderate reasoning, we use powerful MLLMs (Qwen2.5-VL-72B/32B) to generate explicit, step-by-step explanations, creating a corpus of ~12.2 million short CoT samples.
  • Long CoT Enrichment Loop: For the most complex instructions, we leverage top proprietary MLLMs to generate detailed, multi-step solutions, yielding a high-quality set of ~2.7 million long CoT pairs.
  • Fidelity Verification: Throughout the enrichment process, a verifier model (LLM-as-a-Judge) performs semantic comparisons to ensure the correctness and consistency of the generated CoT responses.

🍯 Honey-Data-15M

The primary output of our pipeline is Honey-Data-15M, a large-scale, multimodal SFT dataset with 15 million meticulously curated samples. It is designed to serve as a new cornerstone for the fully open MLLM community. A defining feature is its enrichment with dual-level CoT reasoningβ€”approximately 12.2 million short CoT samples and 2.7 million long CoT samplesβ€”which provides tailored reasoning depth across a wide spectrum of critical domains like "General" visual understanding and "STEM" for symbolic reasoning.

Honey-Data-15M Category Distribution

Figure 2: Category distribution of Honey-Data-15M dataset.

Honey-Data-15M Distribution Pie Chart

Figure 3: Data collection of Honey-Data-15M. A detailed breakdown of our dataset's composition across seven major categories. The number of samples (in thousands) is listed for each source. The * denotes that the data contains the long CoT response.

🍯 Honey-Data-V2 New

With the journal version we release Honey-Data-V2, produced by reapplying HoneyPipe to Honey-Data-15M at a substantially larger scale: 44.3M samples comprising 100.3M QA pairs over 19.7M image groups (5.81 TB), with long CoT covering 17.9% of the corpus (7.9M samples, against 2.9M in Honey-Data-15M). It is not built from scratch; it extends Honey-Data-15M in three ways, all executed with the existing operators of DataStudio:

  • Re-curation of the existing pool. Acting on community defect reports, samples whose final turn carries no answer were removed, 492.6K grounding instructions were rewritten to state their coordinate convention up front, and 1.8M machine-recaptioning samples were dropped. Image placeholders were canonicalized, and degenerate generations (unpaired reasoning tags, repetition collapse, placeholder leakage) were rejected.
  • Multi-model rollout. On top of the original annotation pass, two further independent rollouts answered the same instructions with an upgraded model stack (Qwen3-VL-235B-A22B Instruct and Thinking, Doubao-Thinking, Kimi-K2.5), all checked by the same verifier, Qwen3-235B-A22B-Instruct. Every verified answer is kept, so an instruction carries one to three independently produced answers.
  • New source incorporation. The FineVision datasets absent from Honey-Data-15M, together with Molmo2-SynMultiImageQA, ChartVerse and Aria-UI, were passed through the same HoneyPipe rather than concatenated as-is. Sources whose answers are already determinate, such as OCR transcription and box-level grounding, were filtered but not rewritten.
Lineage Samples Long CoT
Re-curated Honey-Data-15M pool (three rollout passes: 9.6M + 12.8M + 7.0M)29.4M15.5%
Newly incorporated sources14.9M22.6%
Honey-Data-V2 total44.3M17.9%
Honey-Data-V2 composition by domain

Figure 4: Composition of Honey-Data-V2 by domain and CoT level. Darker segments denote long CoT.

Honey-Data-V2 sources of each domain

Figure 5: Sources of each domain: the three rollout passes over the Honey-Data-15M pool versus the newly incorporated sources.

Each record of the release groups every conversation that refers to the same image, so image bytes are stored once, and every answer carries the provenance of the models that wrote and verified it. See the dataset card for the record layout and loading examples. The Bee-8B results below are from models trained on Honey-Data-15M; evaluating Honey-Data-V2 is left for future work.

🎯 Model-Tailored RL Data

The RL stage refines the reasoning patterns instilled by SFT, mitigates formatting errors, and encourages rigorous logical deduction. Its data is curated to match the model's current capacity: open-source RL datasets and grounding and counting tasks transformed from our SFT grounding sources form the initial pool, and the SFT model then generates eight rollouts per prompt. An LLM judge scores them, and prompts the model already solves in all eight rollouts are discarded, leaving only appropriately challenging samples with definitive ground truths for GRPO training.

Model-tailored RL data curation pipeline

Figure 6: Overview of the model-tailored RL data curation pipeline. The data aggregation stage merges open-source and grounding data into an initial pool. During capability-tailored filtering, the SFT model generates 8 rollouts per prompt. A judge model evaluates them, discarding overly easy tasks (all 8 correct) to retain only appropriately challenging samples.

🐝 Overall results for Bee-8B

To validate our Honey-Data-15M, we trained Bee-8B, a new 8B parameter model, based on Qwen3-8B, on the full Honey-Data-15M dataset. Bee-8B establishes a new performance bar for fully open models, particularly in factual accuracy and complex reasoning, and proves highly competitive with recent semi-open models. These results confirm our core thesis: a focus on high-quality data curation is critical for creating models that can rival leading semi-open counterparts. The table reports the released Bee-8B-SFT and Bee-8B-RL checkpoints.

Task Benchmark LLaVA
OneVision-7B*
Molmo
-7B-D*
Qwen2.5
-VL-7B†
Keye-VL
-8B†
InternVL3.5
-8B†
Bee-8B
-SFT*
Bee-8B
-RL*
General
VQA
AI2D 81.4 81.0 84.3 86.7 84.0 83.8 85.3
BLINKval 48.2 49.7 56.4 52.0 59.5 52.5 55.0
CountBench β€” 84.8 74.1 78.0 β€” 90.5 93.0
HallusionBenchavg 31.6 46.4 52.9 67.0 54.5 59.8 58.2
MMBench-CNdev β€” β€” 81.3 92.0 β€” 81.2 84.2
MMBench-ENdev 80.8 β€” 82.1 91.5 β€” 83.0 85.5
MMMUval 48.8 45.3 58.6 71.4 73.4 66.8 66.1
MMMU-Prostandard 29.5 β€” 34.7 47.1 β€” 50.4 50.7
MMStar 61.7 56.1 63.9 75.5 69.3 69.0 71.4
MMT-Benchval 59.3 56.3 63.6 65.9 66.7 64.6 67.0
MMVet 57.5 41.5 67.1 79.0 83.1 83.3 83.9
MMVP β€” β€” 73.3 79.0 β€” 80.7 82.0
POPEavg 88.4 89.0 86.4 86.0 88.7 84.0 84.8
RealWorldQA 66.3 70.7 68.5 67.7 67.5 70.1 73.1
VisuLogic β€” β€” 20.0 25.6 β€” 24.4 26.5
VLMs are Blind 39.2 β€” 37.4 57.1 β€” 55.8 56.5
Table & Chart
& OCR
CharXivDQ β€” β€” 73.9 77.7 72.2 84.7 84.8
CharXivRQ β€” β€” 42.5 45.4 44.4 55.3 57.3
ChartQAtest 80.0 84.1 87.3 86.3 86.7 86.7 86.1
DocVQAval β€” β€” 95.5 88.5 β€” 87.2 87.0
InfoVQAval β€” β€” 81.4 67.4 β€” 72.3 72.9
OCRBench 62.2 65.6 86.4 85.1 84.0 83.1 82.5
SEED-Bench2-Plus 65.4 67.6 70.4 69.4 70.8 67.7 68.5
Math &
Reasoning
DynaMathworst 9.0 β€” 21.0 37.3 37.7 41.3 40.5
LogicVista 33.3 β€” 44.1 54.8 57.3 56.8 61.3
MathVersevision_only 26.2 4.2 25.1 59.8 61.5 61.9 67.0
MathVision 18.3 16.2 25.4 46.0 56.8 46.8 50.0
MathVistamini 63.2 51.6 68.2 80.7 78.4 78.6 81.4
WeMath 20.9 β€” 35.2 60.7 57.0 55.0 59.8

Table 1: Performance comparison of Bee-8B with other fully open (*) and semi-open (†) models across various benchmarks. The top and second-best scores for each benchmark are highlighted.

BibTeX

Journal version (Honey-Data-V2):

@article{zhang2026beecvm,
  title={Bee: A High-Quality Corpus and Full-Stack Suite for End-to-End Training of Fully Open MLLMs},
  author={Zhang, Yi and Ni, Bolin and Chen, Xin-Sheng and Zhang, Heng-Rui and Rao, Yongming and Peng, Houwen and Lu, Qinglin and Hu, Han and Guo, Meng-Hao and Hu, Shi-Min},
  journal={Computational Visual Media},
  year={2026},
  note={to appear}
}

Conference version (Honey-Data-15M, Bee-8B):

@inproceedings{zhang2026bee,
  title={Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs},
  author={Zhang, Yi and Ni, Bolin and Chen, Xin-Sheng and Zhang, Heng-Rui and Rao, Yongming and Peng, Houwen and Lu, Qinglin and Hu, Han and Guo, Meng-Hao and Hu, Shi-Min},
  booktitle={International Conference on Learning Representations (ICLR)},
  year={2026}
}