基于预训练二维模型和扩散模型的点云理解方法

DIFF3DU: PRE-TRAINED 2D DIFFUSION MODELS AS 3D POINT CLOUD UNDERSTANERS

  • 摘要: 近年来,三维视觉领域受到了广泛关注,而三维标注数据的获取,仍是一个巨大的挑战,阻碍了大规模预训练模型在三维视觉领域的发展。因此,将三维点云投影为二维图像,从而将二维预训练模型中的知识用于三维点云理解被认为是一种十分有前景的方法。然而,缩小三维和二维之间的领域差距并消除三维到二维投影过程中引入的噪声仍是一项挑战。针对这一问题,提出一种名为Diff3DU的方法,该方法利用预训练的扩散模型来缓解噪声问题,并提高二维预训练模型在三维点云理解上的性能。将实验结果与其他方法进行综合比较和分析,证明了该方法的有效性。

     

    Abstract: In recent years, the field of 3D vision has garnered substantial attention. However, acquiring 3D data, particularly annotated data, poses a significant challenge that impedes the progress of large- scale pre- trained models in the realm of 3D vision. In this context, the inherent abundant knowledge present in 2D pre- trained models presents a promising approach for understanding 3D point clouds by directly projecting them onto 2D images, demonstrating both feasibility and potential. However, it is still challenging to narrow the domain gap between 3D and 2D and eliminate the noise introduced during the 3D- to- 2D projection process. Therefore, this paper proposes a novel method dubbed Diff3DU that utilizes a pre- trained diffusion model to alleviate the noise problem and improve the performance of the 2D pre- trained models on 3D point cloud understanding. Comprehensive comparison and analysis of the experimental results against widely- used benchmarks across diverse settings demonstrate the effectiveness of the proposed method.

     

/

返回文章
返回