Hierarchical Text-Conditional Image Generation with CLIP Latents
Proposes a two-stage framework using CLIP latents and diffusion models to enhance image diversity and zero-shot control.
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol et al.
Proposes a two-stage framework using CLIP latents and diffusion models to enhance image diversity and zero-shot control.
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol et al.
ASQA dataset addresses ambiguity in long-form QA with a new evaluation metric.
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra et al.
TopFormer employs a multi-scale Token Pyramid and scale-aware semantics, achieving 5% higher mIoU with low latency on mobile devices.
Wenqiang Zhang, Zilong Huang, Guozhong Luo et al.
Debate-style explanations trained with long contexts did not significantly improve human answer accuracy; snippets outperformed debates.
Alicia Parrish, Harsh Trivedi, Ethan Perez et al.
Proposes a linear-time algorithm for Bernstein-Bézier coefficients of B-spline basis functions, optimizing computational complexity.
Filip Chudy, Paweł Woźny
Proposed a hierarchical OPF algorithm with improved gradient evaluation, enhancing voltage safety and computational efficiency.
Heng Liang, Xinyang Zhou, Changhong Zhao
NAFNet eliminates nonlinear activations, surpasses SOTA with 33.69dB PSNR on GoPro, using only 8.4% of the original computational cost.
Liangyu Chen, Xiaojie Chu, Xiangyu Zhang et al.
LaF: Bayesian-based, labeling-free model ranking method, improves performance estimation without labels.
Qiang Hu, Yuejun Guo, Maxime Cordy et al.
Pure tactile in-hand manipulation using torque-controlled hand trained with SAC in 600 CPU hours, achieving over 46 rotations.
Leon Sievers, Johannes Pitz, Berthold Bäuml
UniCL unifies contrastive learning in image-text-label space, achieving 14.5% improvement in zero-shot benchmarks.
Jianwei Yang, Chunyuan Li, Pengchuan Zhang et al.
Video diffusion model using 3D U-Net architecture enables high-quality long video synthesis with conditional sampling and joint training.
Jonathan Ho, Tim Salimans, Alexey Gritsenko et al.
Open-source benchmark framework for VG, analyzing impact of components on recall@N, model size, and efficiency.
Gabriele Berton, Riccardo Mereu, Gabriele Trivigno et al.
Survey of 279 UX practitioners reveals collaboration challenges in data merging, mainly resource shortages and disagreements, impacting analysis reliability.
Emily Kuang, Xiaofu Jin, Mingming Fan
ObjectFolder 2.0 uses implicit neural representations to build a large-scale multisensory dataset, enabling effective Sim2Real transfer for robotics and vision tasks.
Ruohan Gao, Zilin Si, Yen-Yu Chang et al.
Pathways系统支持下的540B参数PaLM模型,显著提升少样学习能力,超越多项自然语言任务的SOTA,展现出大规模模型的潜力。
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin et al.
Proposed an autoregressive 3D shape generation method using Transformers, achieving efficient generation via semantically aligned shape composition sequences.
An-Chieh Cheng, Xueting Li, Sifei Liu et al.
This survey systematically reviews graph embedding methods, including traditional and GNN-based techniques for static and dynamic graphs, analyzing over 300 papers since 2017.
Shima Khoshraftar, Aijun An
This study evaluates Nomon, a probabilistic interface for noisy single-switch users, showing it outperforms row-column scanning in long-term tasks with specific metrics.
Nicholas Bonaker, Emli-Mari Nel, Keith Vertanen et al.
ESCM$^2$ model uses counterfactual risk minimization to address CVR estimation bias, significantly enhancing recommender system performance.
Hao Wang, Tai-Wei Chang, Tianqiao Liu et al.
Proposes a physics-guided neural operator learning approach to predict biological tissue displacement fields from DIC measurements, outperforming traditional models.
Huaiqian You, Quinn Zhang, Colton J. Ross et al.