Submitted by
chenkangjie1123
XPENG AI
non-profit
AI & ML interests
LLM, VLM, Omni Model, Agent, VLA
Recent Activity
View all activity
Papers
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation