cs.CV 2605.21171

FTerViT: Fully Ternary Vision Transformer

FTerViT fully ternarizes all weights and normalization parameters, achieving 82.43% accuracy with 15× compression on ImageNet.

Szymon Ruciński, Pietro Bonazzi, Engin Türetken et al.

2026-05-20 35
cs.CV 2605.21061

Grounding Driving VLA via Inverse Kinematics

Transforming driving VLA into an inverse kinematics framework with future visual prediction and diffusion-based IK network, achieving performance comparable to much larger models.

Junsung Park, Hyunjung Shim

2026-05-20 43
cs.CV 2605.19804

Stitched Value Model for Diffusion Alignment

StitchVM efficiently transfers pretrained pixel reward models into noisy latent space, boosting speed and robustness.

Hyojun Go, Hyungjin Chung, Prune Truong et al.

2026-05-19 44