cs.CV 2605.21171

FTerViT: Fully Ternary Vision Transformer

FTerViT fully ternarizes all weights and normalization parameters, achieving 82.43% accuracy with 15× compression on ImageNet.

Szymon Ruciński, Pietro Bonazzi, Engin Türetken et al.

2026-05-20 41
cs.CV 2605.21061

Grounding Driving VLA via Inverse Kinematics

Transforming driving VLA into an inverse kinematics framework with future visual prediction and diffusion-based IK network, achieving performance comparable to much larger models.

Junsung Park, Hyunjung Shim

2026-05-20 49