项目概述

UniDepthV2 由 Luigi Piccinelli 等人在 2025 年 2 月发布。 UniDepthV2，能够跨域仅从单张图像重建度量三维场景。与现有的 MMDE 范式不同，UniDepthV2 在推理时直接从输入图像预测度量三维点，无需任何额外信息，力求实现通用且灵活的 MMDE 解决方案。相关论文成果为 UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler 。

本教程采用资源为单卡 RTX 4090 。

引用信息

本项目引用信息如下：

@inproceedings{piccinelli2024unidepth, title = { {U}ni{D}epth: Universal Monocular Metric Depth Estimation}, author = {Piccinelli, Luigi and Yang, Yung-Hsu and Sakaridis, Christos and Segu, Mattia and Li, Siyuan and Van Gool, Luc and Yu, Fisher}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2024} } @misc{piccinelli2025unidepthv2, title={ {U}ni{D}epth{V2}: Universal Monocular Metric Depth Estimation Made Simpler}, author={Luigi Piccinelli and Christos Sakaridis and Yung-Hsu Yang and Mattia Segu and Siyuan Li and Wim Abbeloos and Luc Van Gool}, year={2025}, eprint={2502.20110}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2502.20110}, }

HyperAI

运行此教程在 Discord 上讨论

日期

8 个月前

大小

722.73 MB

标签

深度估计

计算机视觉

许可证

Other

GitHub

lpiccinelli-eth/UniDepth

论文 URL

2502.20110

项目概述

本教程采用资源为单卡 RTX 4090 。

项目示例

运行步骤

1. 启动容器后点击 API 地址即可进入 Web 界面

若显示「Bad Gateway」，这表示模型正在初始化，由于模型较大，请等待约 1-2 分钟后刷新页面。

2. 进入网页后，即可与模型进行交互

交流探讨

🖌️ 如果大家看到优质项目，欢迎后台留言推荐！另外，我们还建立了教程交流群，欢迎小伙伴们扫码备注【SD 教程】入群探讨各类技术问题、分享应用效果↓

引用信息

本项目引用信息如下：

@inproceedings{piccinelli2024unidepth,
    title     = { {U}ni{D}epth: Universal Monocular Metric Depth Estimation},
    author    = {Piccinelli, Luigi and Yang, Yung-Hsu and Sakaridis, Christos and Segu, Mattia and Li, Siyuan and Van Gool, Luc and Yu, Fisher},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year      = {2024}
}

@misc{piccinelli2025unidepthv2,
      title={ {U}ni{D}epth{V2}: Universal Monocular Metric Depth Estimation Made Simpler}, 
      author={Luigi Piccinelli and Christos Sakaridis and Yung-Hsu Yang and Mattia Segu and Siyuan Li and Wim Abbeloos and Luc Van Gool},
      year={2025},
      eprint={2502.20110},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2502.20110}, 
}

该教程由社区用户贡献，仅供交流学习使用。如内容涉及侵权，请联系邮箱 [email protected] 以便及时审查和下架。

用 AI 构建 AI

从创意到上线——通过免费 AI 协同编码、开箱即用的环境和最优惠的 GPU 价格,加速您的 AI 开发。

AI 协同编码

开箱即用的 GPU

最优定价

开始使用查看定价

HyperAI Newsletters

订阅我们的最新资讯

我们会在北京时间 每周一的上午九点 向您的邮箱投递本周内的最新更新

邮件发送服务由 MailChimp 提供

HyperAI

运行此教程在 Discord 上讨论

日期

8 个月前

大小

722.73 MB

标签

深度估计

计算机视觉

许可证

Other

GitHub

lpiccinelli-eth/UniDepth

论文 URL

2502.20110

项目概述

本教程采用资源为单卡 RTX 4090 。

项目示例

运行步骤

1. 启动容器后点击 API 地址即可进入 Web 界面

若显示「Bad Gateway」，这表示模型正在初始化，由于模型较大，请等待约 1-2 分钟后刷新页面。

2. 进入网页后，即可与模型进行交互

交流探讨

引用信息

本项目引用信息如下：

@inproceedings{piccinelli2024unidepth,
    title     = { {U}ni{D}epth: Universal Monocular Metric Depth Estimation},
    author    = {Piccinelli, Luigi and Yang, Yung-Hsu and Sakaridis, Christos and Segu, Mattia and Li, Siyuan and Van Gool, Luc and Yu, Fisher},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year      = {2024}
}

@misc{piccinelli2025unidepthv2,
      title={ {U}ni{D}epth{V2}: Universal Monocular Metric Depth Estimation Made Simpler}, 
      author={Luigi Piccinelli and Christos Sakaridis and Yung-Hsu Yang and Mattia Segu and Siyuan Li and Wim Abbeloos and Luc Van Gool},
      year={2025},
      eprint={2502.20110},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2502.20110}, 
}

该教程由社区用户贡献，仅供交流学习使用。如内容涉及侵权，请联系邮箱 [email protected] 以便及时审查和下架。

Depth-Anything-3：从任何视角恢复视觉空间

2 个月前

HunyuanWorld-Mirror：3D 世界生成模型

2 个月前

HunyuanOCR：腾讯混元端到端 OCR

2 个月前

Supertonic：基于 ONNX 的极速 TTS 语音合成模型

2 个月前

腾讯混元 HunyuanVideo-Foley

1 个月前

DiffVox：声音区分效果模型

2 个月前

VibeVoice-Realtime TTS：实时语音合成服务

2 个月前

Open-AutoGLM：手机端智能助理

2 个月前

Kiss3DGen：基于图像扩散模型的 3D 资产生成框架

1 个月前

用 AI 构建 AI

从创意到上线——通过免费 AI 协同编码、开箱即用的环境和最优惠的 GPU 价格,加速您的 AI 开发。

AI 协同编码

开箱即用的 GPU

最优定价

开始使用查看定价

HyperAI Newsletters

订阅我们的最新资讯

我们会在北京时间 每周一的上午九点 向您的邮箱投递本周内的最新更新

邮件发送服务由 MailChimp 提供

Command Palette

UniDepthV2：通用单目度量深度估计

项目概述

项目示例

运行步骤

交流探讨

引用信息

用 AI 构建 AI

HyperAI Newsletters

Command Palette

UniDepthV2：通用单目度量深度估计

项目概述

项目示例

运行步骤

交流探讨

引用信息

相关教程

Depth-Anything-3：从任何视角恢复视觉空间

HunyuanWorld-Mirror：3D 世界生成模型

HunyuanOCR：腾讯混元端到端 OCR

Supertonic：基于 ONNX 的极速 TTS 语音合成模型

腾讯混元 HunyuanVideo-Foley

DiffVox：声音区分效果模型

VibeVoice-Realtime TTS：实时语音合成服务

Open-AutoGLM：手机端智能助理

Kiss3DGen：基于图像扩散模型的 3D 资产生成框架

用 AI 构建 AI

HyperAI Newsletters

Command Palette

UniDepthV2：通用单目度量深度估计

项目概述

项目示例

运行步骤

交流探讨

引用信息

相关教程

Depth-Anything-3：从任何视角恢复视觉空间

HunyuanWorld-Mirror：3D 世界生成模型

HunyuanOCR：腾讯混元端到端 OCR

Supertonic：基于 ONNX 的极速 TTS 语音合成模型

腾讯混元 HunyuanVideo-Foley

DiffVox：声音区分效果模型

VibeVoice-Realtime TTS：实时语音合成服务

Open-AutoGLM：手机端智能助理

Kiss3DGen：基于图像扩散模型的 3D 资产生成框架

用 AI 构建 AI

HyperAI Newsletters

相关教程

Depth-Anything-3：从任何视角恢复视觉空间

HunyuanWorld-Mirror：3D 世界生成模型

HunyuanOCR：腾讯混元端到端 OCR

Supertonic：基于 ONNX 的极速 TTS 语音合成模型

腾讯混元 HunyuanVideo-Foley

DiffVox：声音区分效果模型

VibeVoice-Realtime TTS：实时语音合成服务

Open-AutoGLM：手机端智能助理

Kiss3DGen：基于图像扩散模型的 3D 资产生成框架

相关教程

Depth-Anything-3：从任何视角恢复视觉空间

HunyuanWorld-Mirror：3D 世界生成模型

HunyuanOCR：腾讯混元端到端 OCR

Supertonic：基于 ONNX 的极速 TTS 语音合成模型

腾讯混元 HunyuanVideo-Foley

DiffVox：声音区分效果模型

VibeVoice-Realtime TTS：实时语音合成服务

Open-AutoGLM：手机端智能助理

Kiss3DGen：基于图像扩散模型的 3D 资产生成框架