RGB-depth cameras play a critical role in applications such as 3D reconstruction, scanning, and Simultaneous Localization and Mapping (SLAM), which must undergo rigorous testing to ensure optimal performance. While qualitative testing can confirm basic functionality, it is the quantitative testing that truly matters for applications requiring precise depth readings and measurements.
In this blog, we encapsulate the key insights Xin shared at the Vision Spectra Conference 2024 on the methods and metrics used to evaluate depth cameras, with a focus on both qualitative and quantitative testing approaches, offering an in-depth exploration of the performance metrics associated with various depth-sensing technologies.
A Deeper Dive into Quantitative Testing
Quantitative testing is essential for capturing the numerical data required for high-accuracy applications. Xin highlighted the three primary depth-sensing technologies available today:
- Structured Light Depth Cameras
- Stereo Vision Depth Cameras
- Time of Flight (TOF) Depth Cameras
Each of these technologies operates on unique principles to measure distance. For instance, structured light cameras calculate depth by comparing a captured image with a pre-recorded calibration image, while TOF cameras measure depth based on the time it takes for a laser to bounce back from an object.
Accuracy and Error Analysis
To accurately assess these cameras, we have developed a comprehensive testing workflow. This involves:
- Setting up a testing environment with a white flat wall as the target.
- Using an optical rail to ensure the camera and laser distance meter move rigidly together.
- Capturing data from varying distances (0.5 meters to 6 meters) at intervals of 100 millimeters.
- Turning off all post-processing filters to ensure raw data comparison.

We selected multiple regions of interest (ROI) within the captured depth images to evaluate the presence of lens distortion and other anomalies. This multi-ROI approach helps us understand if the camera’s performance varies across different areas of the frame.

Data Presentation and Analysis
The analysis of the captured data involves several key metrics:
- Accuracy: Calculated by comparing depth readings to ground truth values obtained from a laser distance meter.
- Root Mean Square Error (RMSE): A widely used method to represent error.
- Standard Deviation (STD): Indicates the stability and concentration of data.

We visualize this data using charts and histograms, which provide a clear picture of each camera’s performance. For instance, our findings show that TOF cameras generally offer the best accuracy, while structured light cameras perform well at close ranges.
Exploring Spatial and Temporal Precision Metrics
Spatial precision refers to the consistency of depth readings across different parts of the image frame. We calculate the mean STD within selected ROIs and plot these values against distance. Interestingly, while TOF cameras lead in accuracy, structured light cameras show superior spatial precision at close ranges, followed by stereo vision cameras.
Temporal precision evaluates the stability of depth readings over time. In our tests, we captured depth images every second for ten minutes at a fixed distance. A stable camera should show minimal variation in these readings over time, indicating reliability for long-term applications.
Practical Insights and Comparisons
Our testing framework not only highlights the strengths and limitations of each depth-sensing technology but also provides practical insights for real-world applications. For example, TOF cameras excel in accuracy over longer distances, making them ideal for industrial applications, while structured light cameras are more cost-effective for close-range uses, such as Apple’s Face ID.
Xin also shared specific examples of cameras with performance issues, such as lens distortion leading to significant depth inaccuracies at the edges of the frame. He used heatmaps to demonstrate how TOF cameras showed increased error rates at distances beyond 4 meters, and histograms to illustrate the temporal instability of certain stereo vision cameras over prolonged use. These analyses are crucial for identifying and addressing potential problems in camera design and manufacturing, helping us refine our technologies and improve overall performance.
In conclusion, we are committed to advancing precise and reliable depth-sensing technology. Our rigorous testing methodologies ensure that our cameras meet the highest standards of accuracy and performance, paving the way for innovative applications across various industries.
Watch Xin’s full webinar here.
Frequently Asked Questions About Decoding Depth Camera Performance
1. Why does quantitative testing matter when evaluating a depth camera?
RGB-depth cameras play a critical role in applications such as 3D reconstruction, scanning, and Simultaneous Localization and Mapping (SLAM), all of which require rigorous testing to ensure optimal performance. Qualitative testing can confirm that a camera basically functions, but quantitative testing is what captures the numerical data that high-accuracy applications depend on. The blog summarizes insights Xin shared at the Vision Spectra Conference 2024 on the methods and metrics used to evaluate depth cameras.
2. What testing setup is used to measure depth camera accuracy?
The testing workflow uses a white flat wall as the target and an optical rail so that the camera and the laser distance meter move rigidly together. Data is captured from varying distances between 0.5 meters and 6 meters at intervals of 100 millimeters, with all post-processing filters turned off so that raw data can be compared. Multiple regions of interest are selected within the captured depth images to evaluate lens distortion and other anomalies, and this multi-ROI approach reveals whether a camera’s performance varies across different areas of the frame.
3. Which metrics are used to analyze depth camera performance?
Three key metrics drive the analysis. Accuracy is calculated by comparing depth readings against ground truth values obtained from a laser distance meter. Root Mean Square Error, or RMSE, is a widely used method of representing error. Standard Deviation, or STD, indicates the stability and concentration of the data. The results are visualized using charts and histograms to give a clear picture of each camera’s performance.
4. What is the difference between spatial precision and temporal precision?
Spatial precision refers to the consistency of depth readings across different parts of the image frame, and it is measured by calculating the mean STD within selected regions of interest and plotting those values against distance. Temporal precision evaluates the stability of depth readings over time. In these tests, depth images were captured every second for ten minutes at a fixed distance, since a stable camera should show minimal variation across those readings, which indicates reliability for long-term applications.
5. How do structured light, stereo vision, and TOF cameras compare?
The three primary depth-sensing technologies each operate on different principles: structured light cameras calculate depth by comparing a captured image with a pre-recorded calibration image, while TOF cameras measure depth based on the time it takes a laser to bounce back from an object. In testing, TOF cameras generally offered the best accuracy, while structured light cameras performed well at close ranges. On spatial precision the picture shifts, with structured light cameras showing superior spatial precision at close ranges, followed by stereo vision cameras. In practical terms, TOF cameras excel in accuracy over longer distances, which suits industrial applications, while structured light cameras are more cost-effective for close-range uses such as Apple’s Face ID.
6. What performance problems does this kind of testing reveal?
The testing surfaced several concrete issues. Lens distortion in some cameras led to significant depth inaccuracies at the edges of the frame. Heatmaps showed TOF cameras with increased error rates at distances beyond 4 meters. Histograms illustrated the temporal instability of certain stereo vision cameras over prolonged use. Identifying problems like these is important for addressing weaknesses in camera design and manufacturing and for refining the technology to improve overall performance.



