DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

The University of Hong Kong
* Project Leader. † Corresponding author.
DVD teaser

Comparison of MLLM perception paradigms. Existing methods either use text tokens (high token count, e.g., 512 tokens for masks) or special coordinate tokens (limited range and precision) to encode perception outputs. DVD unifies diverse 2D/3D representations into 1D sequences with a unified codebook in high-dimensional space, yielding compact and high-precision representations with fewer tokens and consistent latency reduction across 2D and 3D tasks.

Existing MLLM-based perception methods rely on text-based coordinates or fixed-range special tokens, suffering from excessive token overhead and precision constraints, especially in 3D domains. DVD (Dynamic Vector Decoding) converts 2D boxes, 2D masks, and 3D boxes into 1D vector sequences, and learns a unified codebook in high-dimensional space. A lightweight de-tokenizer then maps the discrete tokens back to original perceptual representations, enabling seamless integration with MLLMs while dramatically reducing token count and inference latency.

Abstract

Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representations (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, and nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and significantly reduces token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.

DVD Framework

Why Existing MLLM Perception is Inefficient

Existing MLLM-based perception methods typically represent coordinates as text tokens or discrete special tokens within a fixed integer range. Text-based methods can represent arbitrary precision and range, but they generate a large number of tokens (e.g., 512 tokens for a mask) and force the model to learn the structure of floating-point numbers. Fixed-range quantization reduces token count, but it sacrifices range and precision—critical limitations when moving from constrained 2D images to unbounded 3D scenes.

Unified 1D Vector Representation

DVD unifies heterogeneous perceptual outputs by flattening them into 1D vector sequences:

  • 2D bounding box: $(x_1, y_1, x_2, y_2)$
  • 2D mask: fixed number of edge points $(x_1, y_1, \ldots, x_N, y_N)$ sampled clockwise and normalized by image size
  • 3D bounding box: $(x, y, z, w, h, l, R)$, where $R \in \mathbb{R}^{3 \times 3}$ is the rotation matrix; raw values are preserved to maintain unbounded range and precision

Dynamic Vector Representation Learning

Dynamic vector representation learning

Dynamic vector representation learning. 1D sequences serve as unified 2D/3D perceptual representations; a 1D encoder maps them into high-dimensional space for adaptive quantization and compression, a task-agnostic geometric loss supervises the heterogeneous representation learning, and a decoder recovers the original representations from the discrete tokens.

A 1D encoder-decoder architecture maps the vector sequences into a latent space and reconstructs them via a learnable codebook. Given an input sequence $\mathbf{s} \in \mathbb{R}^{B \times L \times 1}$, the encoder produces latent vectors $\mathbf{z} \in \mathbb{R}^{B \times \frac{L}{K} \times C}$, where $K$ is the downsampling factor and $C$ is the feature dimension. Each latent vector is quantized to the nearest codebook entry. The symmetric decoder upsamples the quantized vectors back to the original sequence.

Beyond the standard reconstruction and VQ losses, we introduce a task-agnostic geometric loss based on the bidirectional chamfer distance. The per-dimension L1 loss is sensitive to point ordering and parameterization rather than the underlying geometry, while the VQ loss only constrains the latent space. The geometric loss is permutation-invariant: for 3D tasks it computes the average minimum L1 distance between predicted and target 3D box vertices, and for 2D tasks between predicted and target 2D point sets, providing unified geometric supervision for both 2D and 3D perception tasks. The model is trained end-to-end with a combined loss:

$\mathcal{L}_{\text{total}} = \lambda_1 \cdot \mathcal{L}_{\text{L1}} + \lambda_2 \cdot \mathcal{L}_{\text{VQ}} + \lambda_3 \cdot \mathcal{L}_{\text{geo}}$

where the L1 loss ensures reconstruction accuracy, the VQ loss updates the codebook, and the geometric loss preserves the spatial structure (corners, edges, and shapes) of the decoded representations.

Dynamic Vector Decoding with MLLMs

DVD architecture overview

Overview of DVD. The MLLM vocabulary is extended with learned unified tokens and trained following standard LLM paradigms without modifications; at inference, a lightweight de-tokenizer reconstructs the original perceptual outputs from the discrete tokens with minimal overhead, drastically reducing sequence length and improving efficiency.

To integrate with MLLMs, we extend the vocabulary with $M$ new tokens corresponding to the $M$ code vectors in the codebook. The MLLM is then fine-tuned with the standard next-token prediction objective, without modifying the backbone. During inference, perception tokens from the MLLM output are fed into the pre-trained decoder (used as a de-tokenizer) to recover the 1D vector sequence, which is then reshaped into 2D boxes, masks, or 3D boxes. Because each perceptual output is represented by only a few tokens, the sequence length fed to the MLLM is greatly reduced, cutting latency while preserving spatial precision.

Experiments

Datasets and Metrics

We evaluate DVD on 2D and 3D perception benchmarks. For 3D grounding, we use SUN-RGBD, Hypersim, ARKitScenes, KITTI, and nuScenes, covering indoor and outdoor scenarios, and report AP3D@15 following Omni3D. For 2D grounding and referring expression segmentation (RES), we use RefCOCO, RefCOCO+, and RefCOCOg, and report Precision@0.5 and cIoU, respectively. The same trained model is used to validate the performance of all tasks.

The unified vector autoencoder is trained on approximately 1.98M 3D boxes, 2.31M 2D boxes, and 1.52M 2D masks, with downsampling factor $K{=}2$, codebook size 4096, and latent channel 16. The MLLM is built on Qwen3-VL (2B and 8B variants) and fully fine-tuned on 2.0M mixed 2D/3D perception QAs.

Overall Performance

Benefiting from dynamic vector decoding, DVD-2B surpasses SOTA MLLM-based methods on 3D grounding with 40.0%, 15.7%, and 19.8% AP3D on SUN-RGBD, Hypersim, and nuScenes, and achieves 89.9% Precision and 70.4% cIoU on RefCOCOg val — all with a single unified 2B model. With a larger backbone, DVD-8B further reaches 40.3%, 18.4%, and 24.4% AP3D on the three 3D datasets, and 88.7% Precision and 73.4% cIoU on RefCOCOg val.

Method 3D Grounding (AP3D@15) 2D Grounding (P@0.5) 2D RES (cIoU)
SUN-RGBD Hypersim nuScenes RefCOCO RefCOCO+ RefCOCOg RefCOCO RefCOCO+ RefCOCOg
GroundingDINO------90.688.286.1------
VistaLLM-7B------88.182.983.674.569.169.0
LISA-7B------------74.965.167.9
Text4Seg------88.383.582.474.768.570.7
Qwen2.5-VL-3B------89.182.485.2------
Qwen2.5-VL-7B------90.084.287.2------
Rex-Omni----------86.6------
Seed1.5-VL33.5----------------
Qwen3-VL-2B33.812.0--89.280.484.0------
Qwen3-VL-8B36.212.7--91.686.187.7------
VST-3B37.3----------------
DVD-2B40.015.719.893.587.789.969.964.270.4
DVD-8B40.318.424.493.087.788.775.470.173.4

Main results on 3D grounding (AP3D@15), 2D grounding (P@0.5), and 2D RES (cIoU). A single unified model is used for all tasks.

3D Grounding

DVD-2B outperforms prior MLLM-based 3D perception methods across indoor and outdoor scenarios, surpassing VST-3B by 2.7% and 10.7% AP3D on SUN-RGBD and ARKitScenes. With a larger backbone, DVD-8B further achieves 40.3% and 24.4% AP3D on SUN-RGBD and nuScenes.

Method SUN-RGBD Hypersim ARKitScenes KITTI nuScenes
Gemini 2.0 Pro32.5--------
Gemini 2.5 Pro29.7--------
Seed1.5-VL33.5--------
Qwen3-VL-2B33.812.0------
Qwen3-VL-8B36.212.7------
VST-3B37.3--51.7----
DVD-2B40.015.762.431.419.8
DVD-8B40.318.462.128.724.4

2D Grounding

Despite being designed for unified 2D and 3D perception, DVD still outperforms advanced 2D grounding methods on the RefCOCO series: DVD-2B achieves 90.0% Precision on RefCOCOg test, surpassing Rex-Omni by 3.2%.

Method RefCOCO val RefCOCO testA RefCOCO testB RefCOCO+ val RefCOCO+ testA RefCOCO+ testB RefCOCOg val RefCOCOg test
GroundingDINO90.693.288.288.289.075.986.187.0
Qwen2.5-VL-7B90.092.585.484.289.176.987.287.2
Rex-Omni------------86.686.8
Seed1.5-VL------------84.785.2
InternVL3.5-8B92.494.788.787.992.482.489.689.4
UFO-8B91.894.387.586.991.380.687.988.6
Qwen3-VL-8B91.693.188.786.191.180.087.788.4
DVD-2B93.595.590.787.792.982.489.990.0
DVD-8B93.095.289.787.792.982.788.789.8

2D Referring Expression Segmentation

DVD also generalizes to fine-grained 2D referring expression segmentation (RES). Despite having far fewer parameters and being a unified model, DVD-2B remains competitive with VistaLLM-7B on RefCOCOg test while running at much lower latency thanks to the reduced token overhead. With a larger backbone, DVD-8B achieves 73.4% and 76.6% cIoU on RefCOCOg val and test.

Method RefCOCO val RefCOCO testA RefCOCO testB RefCOCO+ val RefCOCO+ testA RefCOCO+ testB RefCOCOg val RefCOCOg test
VLT67.570.565.256.361.050.155.057.7
LAVT72.775.868.862.168.455.161.262.1
VistaLLM-7B74.576.072.769.173.764.069.070.9
LISA-7B74.979.172.365.170.858.167.970.6
M2SA74.076.869.763.167.256.167.068.3
PixelLM-7B73.076.568.266.371.758.369.370.5
Text4Seg74.777.471.668.573.662.970.771.6
PerceptionGPT-7B75.178.671.768.573.961.370.371.7
DVD-2B69.970.769.764.268.261.870.471.0
DVD-8B75.476.974.670.174.765.973.476.6

Computation Cost

We compare the average number of tokens generated per object, inference latency, and performance of different representations on SUN-RGBD and RefCOCOg under identical training settings and computation resources. Text-based coordinates generate excessive tokens (512 and 74 for a mask and a 3D box), incurring high latency; while text preserves raw coordinate values and keeps a slight accuracy edge on 3D grounding (37.5% vs. 34.1% AP3D), this marginal gain comes with 9× more tokens and 8.3× higher latency (4165 ms vs. 499 ms per object). Special tokens reduce the cost but degrade severely (17.5% AP3D) due to limited range and precision. In contrast, DVD matches text-level precision on 2D tasks, largely closes the 3D gap, and delivers the best accuracy–efficiency trade-off.

Computation cost and performance comparison of different representations

Computation cost and performance comparison of different representations. Bubble area indicates the token count generated per object of the representation.

Ablation Studies

All ablations are conducted on DVD-2B trained with 30% of the full training iterations for faster validation. We first measure the vector reconstruction quality of the unified vector autoencoder using perception metrics between ground truth and decoded sequences: it achieves negligible reconstruction error — 90.8% AP3D, 100.0% Precision, and 84.8% cIoU on 3D grounding, 2D grounding, and 2D RES — demonstrating its ability to preserve spatial information.

Effect of codebook size

CodebookAP3D@15 avg.P@50 avg.cIoU avg.
204888.984.581.6
409690.8100.084.8
819285.199.676.8

Effect of geometric loss

Geo. LossAP3D@15 avg.P@50 avg.cIoU avg.
✗80.799.174.8
✓90.8100.084.8

Effect of downsampling scale $K$

ScaleAP3D@15 avg.P@50 avg.cIoU avg.
189.3100.092.7
290.8100.084.8
483.731.560.4

Effect of latent hidden size $z$

$z$AP3D@15 avg.P@50 avg.cIoU avg.
883.0100.083.7
1690.8100.084.8
3288.4100.085.3

Codebook size. A small codebook (2048) limits the capacity of the discrete latent space (88.9% AP3D), while an overly large one (8192) makes quantization and codebook learning harder, leaving under-trained entries (85.1% AP3D, 76.8% cIoU). A codebook size of 4096 achieves the best performance across all three tasks and is adopted as the default.

Geometric loss. Geometric guidance consistently improves reconstruction across all tasks, with notable gains on 3D grounding (+10.1% AP3D) and 2D RES (+10.0% cIoU), where fine-grained spatial fidelity (depth estimation, boundary delineation) is sensitive to small coordinate deviations. The gain on 2D grounding is marginal (+0.9% Precision) since coarse box localization is already near saturation — confirming that the geometric loss is an essential component, not merely a regularizer, for unifying 2D/3D grounding and mask-based segmentation in a single codebook.

Downsampling scale. Compared with the uncompressed representation ($K{=}1$), $K{=}2$ halves the token count while maintaining competitive performance and even slightly improving 3D grounding. An excessively large scale ($K{=}4$) causes severe information loss (31.5% Precision on 2D grounding). $K{=}2$ achieves the best trade-off between reconstruction precision and token cost.

Latent hidden size. A narrow bottleneck ($z{=}8$) constrains the encoded spatial information (83.0% AP3D), while enlarging $z$ to 32 brings no consistent gains as higher-dimensional latents are harder to quantize. $z{=}16$ provides sufficient capacity with stable quantization.

3D Grounding 2D Grounding 2D RES AP3D@15 avg. P@50 avg. cIoU avg.
✓84.0----
✓--100.0--
✓----80.3
✓✓✓90.8100.084.8

Reconstruction performance of multi-task training in the unified vector autoencoder.

Multi-task training. Joint multi-task training improves performance on both 3D grounding (84.0% → 90.8% AP3D) and 2D RES (80.3% → 84.8% cIoU) over single-task training, while 2D grounding remains stable without degradation — demonstrating that DVD learns dynamic vector representations without sacrificing individual tasks.

Visualizations

Qualitative Results

The following figures show representative qualitative results of DVD on 3D grounding, 2D grounding, 2D referring segmentation, and unified 2D/3D perception. DVD achieves satisfactory results even in dark and complex 3D scenes, precise 2D localization and segmentation across diverse objects and scenes, and simultaneous 2D and 3D perception within a single model. Click or zoom in to inspect the high-resolution details.

Conclusion

Existing MLLM-based perception methods struggle with excessive token overhead from text-based coordinate representation and precision and range constraints from fixed-range quantization, especially in 3D domains. To address the key limitation, we propose DVD that unifies 2D and 3D perception task representations for efficient integration with MLLMs. By transforming heterogeneous perceptual outputs into 1D vector sequences and mapping them to compact discrete tokens, DVD eliminates the excessive token overhead and the precision limitations. Extensive experiments on 2D and 3D benchmarks show that DVD achieves superior performance while significantly reducing token count and inference latency, offering an efficient and general framework for equipping MLLMs with perception capabilities.

Citation


@article{hou2026dvd,
  title={DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception},
  author={Hou, Jinghua and Liu, Zhe and Zhao, Hengshuang},
  journal={arXiv preprint arXiv:2610.12266},
  year={2026}
}