FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
Abstract
We present Foundation Model Embedded Gaussian Splatting (FMGS), which incorporates vision-language embeddings of foundation models into 3D Gaussian Splatting (GS). The key contribution is an efficient method to reconstruct and represent 3D vision-language models by distilling feature maps generated from image-based foundation models into those rendered from our 3D model. We introduce a novel scene representation by integrating strengths from both GS and multi-resolution hash encodings (MHE), along with a pixel alignment loss that makes the rendered feature distance of the same semantic entities close. Our results demonstrate remarkable multi-view semantic consistency, beating state-of-the-art methods by 10.2 percent on open-vocabulary language-based object detection, despite being 851X faster for inference.
Type
Publication
International Journal of Computer Vision (IJCV)