Search

Hongke's latest articles

HongKe

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Ut elit tellus, luctus nec ullamcorper mattis, pulvinar dapibus leo.

[Hongke Solution] High Fidelity + Strong Control: A Hybrid Rendering Solution for Autonomous Driving Simulation

01 Industry Pain Points and Challenges

Research and development in autonomous driving relies heavily on high-quality data; however, the mainstream simulation testing approaches currently used in the industry each have significant shortcomings:

  • Traditional simulators (such as CARLA): Based on a game engine architecture, scenes offer a high degree of control and objects can be arranged freely; however, the rendering quality lacks realism, resulting in a significant “Domain Shift"—Algorithms that perform well in simulation environments often perform significantly worse when deployed in real vehicles. In addition, building high-precision 3D environments requires a significant amount of manual labor.

  • Neural Reconstruction Methods (NeRF / 3DGS): It can generate photo-realistic renderings, but the scene itself is in a “fixed” state—dynamic objects can only move along their original trajectories. If new objects are inserted or their motion trajectories are altered, the reconstruction quality will deteriorate significantly, making it difficult to support simulation testing for extreme scenarios such as lane-changing to overtake or sudden evasive maneuvers.

Each of these two technical approaches has its own suitable use cases, and deeply integrating their core strengths represents a more practical path to engineering implementation. Building on this approach, this article will explore the application architecture and core value of hybrid rendering in real-world scenarios.

02 How the Two Technologies Can "Complement Each Other"

The basic logic behind blend rendering is as follows:Static environments are rendered using neural reconstruction, while dynamic objects are loaded using traditional graphics technology.The

The overall workflow can be simplified as follows: first, use real-world data to reconstruct static street scenes; then, freely place vehicles and pedestrians within those scenes, while flexibly adjusting the weather and camera angles. Static areas retain the photo-realistic quality of neural rendering, while dynamic objects are rendered using traditional methods to allow for a high degree of control.

Complete Pipeline for Scene Reconstruction

This production line relies on four types of raw input—[in-vehicle camera imagery, LiDAR point clouds, vehicle odometer data, and camera internal and external parameters]—to perform preliminary data processing. Its core comprises four major modules:

  1. Dynamic Object Mask Generation: The solution reconstructs only the static areas of the scene, using OneFormer Panoramic segmentation, optical flow tracking, and LiDAR–camera 3D detection technology identify moving targets such as pedestrians and vehicles; By fusing the results of image-based spatial tracking and 3D bounding box-based static object detection, dynamic masks are generated to mask out moving objects during subsequent model training, ensuring the purity of the static scene reconstruction.

  2. LiDAR Depth and Intensity Map Generation: First, the laser point clouds corresponding to moving objects are removed, and a static aggregated point cloud is obtained through odometer calibration and stitching; then, BEV (Bird's-Eye View) PerspectivePerform Poisson reconstruction in sub-blocks to generate a dense mesh, and then, based on the camera calibration parameters, use Open3D It generates dense depth maps and LiDAR intensity maps, providing precise geometric supervision for neural reconstruction and effectively addressing reconstruction distortion in low-texture areas such as road surfaces.

  3. Camera Pose Precision Optimization: Based on an improved version COLMAP Perform iterative triangulation and bundle adjustment to optimize the camera's internal and external parameters, and then use Kabsch AlgorithmAlign the optimized pose back to the original world coordinate system to eliminate the pose noise inherent in the vehicle's odometer and ensure geometric accuracy in large-scale scene reconstruction.

  4. BEV Block-wise Parallel Processing (BEV): Drawing on SUDS The clustering approach divides the scene into overlapping blocks: the camera coordinates are transformed into BEV space, blocks are partitioned using K-nearest neighbors (KNN) clustering, and the clusters are then extended and expanded using convex hulls to ensure block overlap; The overlapping design eliminates rendering discontinuities (artifacts) at block boundaries, supports cross-block rendering interpolation, adapts to irregular driving trajectories, and enables parallel training of individual blocks, supporting 100,000 square meters or moreReconstruction of ultra-large-scale scenes.

NeRF2GS Two-Stage Reconstruction Core Algorithm

adopt “NeRF Pre-training + 3DGS Real-time Rendering” A two-stage architecture that combines NeRF’s excellent view extrapolation capabilities with the real-time rendering advantages of 3D Gaussian Splatting (3DGS), thoroughly addressing the shortcomings of single-method approaches:

  • Phase 1: Training a Custom NeRF Model

    Based on Nerfstudio's Nerfacto Modify the model to add supervisory constraints across multiple dimensions:

    • Incorporating LiDAR Depth Loss and Sky Region Density $L_2$ Regularization, to suppress floating artifacts on the road surface;

    • A new MLP neural network was developed to predict LiDAR reflectance, and cross-entropy loss was used to perform semantic segmentation prediction across 32 classes;

    • adoptBidirectional-Guided ISP Decoupling Solution, eliminate color inconsistencies caused by the automatic exposure and white balance settings of in-vehicle cameras, and produce a reference scene representation free of imaging artifacts.

  • Phase 2: Multi-Constraint Reinforcement 3DGS Splash Training

    Using the RGB, depth, normal, intensity, and semantic maps output by NeRF as supervision and initialization, we add multiple regularization losses on top of the gsplat backend and the AbsGS refinement strategy:

    • Normal + Flatness Constraint Loss: By aligning the Gaussian distribution with real-world road surfaces and building floor plans, the rendering quality of details such as lane markings and road signs is significantly enhanced from a new perspective.

    • LiDAR Intensity + Semantic Loss: Reflected intensity and 6-bit binary semantic codes are used as Gaussian additional features to enable multimodal information embedding.
    • Gaussian Mechanism for Sky-Ground Separation: Split the scene into a ground Gaussian and a distant spherical sky Gaussian, and use transparency accumulation loss to enforce spatial separation between the two layers, thereby eliminating geometric ambiguity in the background.

Ultimately, it achieved a balance betweenRealistic Geometry, Multimodal Properties, and Real-Time Rendering CapabilitiesStatic scene 3DGS models, serving as the static environment foundation for a hybrid rendering pipeline.

This solution addresses the shortcomings of both types of technologies:Use neural reconstruction to mitigate the “domain shift” in traditional simulation, and use traditional graphics techniques to address the “scene lock-in” limitation of neural reconstruction.The

03 The Three Core Capabilities of a Hybrid Rendering Framework

From the perspective of practical engineering applications, this solution primarily addresses three key issues:Is the scene large enough? Is the image quality detailed enough? Can the generated data be used directly?The

1. The scenes are vast in scale, and the image quality is high and reliable.

Supports block-parallel training 100,000 square meters or morelarge-scale scene reconstruction (roughly equivalent to 14 standard soccer fields). In areas critical to autonomous driving—such as lane markings and road signs—the rendering quality is significantly superior to that of existing, commonly used solutions.

2. Multimodal output, providing comprehensive coverage of mainstream sensors

Can output simultaneously RGB images, depth maps, LiDAR point clouds, surface normal maps, and semantic segmentation masks...a single workflow can comprehensively address the simulation needs of various sensors, including cameras and LiDAR.

3. Efficient operation; data is immediately available

On consumer-grade hardware platforms, the camera rendering speed reaches 83 frames per second (fps), 64-line LiDAR simulation speed reaches 59 frames per second (fps)... and fully supports real-time Hardware-in-the-Loop (HIL) testing.

Data generated using this approach was used to train a 3D object detection model, and the detection accuracy at close range was even higher than when using original real-world images—primarily because the simulated data was more cleanly labeled and had a lower false positive rate. At the same time, the detection results for the simulated images are highly consistent with those for the original images, indicating that domain shift has been minimized. The generated dataCan be used directly for algorithm training, without the need for additional domain adaptation.

04 Conclusion

Discrepancies between simulated environments and real-world road conditions have long hindered the iterative development and deployment of autonomous driving algorithms in actual vehicles. The hybrid rendering approach offers a practical solution—Neural reconstruction preserves realism, while traditional graphics techniques preserve controllability., addressing the two core requirements of realistic scene reproduction and flexible scene editing, while striking the perfect balance between realistic image quality and editing freedom.

The core value of this solution lies in its application of existing technology toHighly Efficient Engineering Integration, effectively addressing practical engineering challenges:

  • For Algorithm Engineers: Acquire large-scale, high-fidelity, and accurately labeled training data at a lower cost;

  • To the Test and Validation Team: Can reproduce various corner cases in a simulation environment to detect potential issues early on;

  • For Project Decision-Makers: Reduce reliance on expensive real-world vehicle test data and significantly accelerate the product iteration cycle.

Essentially, this is a method for generating simulation dataTruly bring the project to fruitionfeasible approaches.

Hongke has long been dedicated to the engineering applications of hybrid rendering technology.aiSim simulation platformDeeply integrating AI rendering technologies such as 3DGS Splash and NeRF, it provides a physically accurate, high-fidelity, and deterministic simulation environment. By combining core capabilities such as high-fidelity scene generation, a library of physically accurate sensor models, and SDG/scene generalization, Kangmou offersA full-process solution ranging from scene construction and synthetic data generation to hardware-in-the-loop (HIL) testing, and is fully committed to enabling the practical application of autonomous driving and robotics technologies.

Other Articles

Hongke Case

[Hongke Solutions] Safety Solution for Long-Distance Ocean Transport of Lithium Batteries – Shengweino International Case Study

How Can Thermal Runaway and Impact Damage Be Prevented During Long-Distance Ocean Transport of Lithium-Ion Batteries? Shengweino International uses Hongke’s ASPION G-Log 2 shock recorder to precisely monitor container impacts (G-forces) and changes in temperature and humidity throughout the entire journey, creating a comprehensive transport environment history and improving the efficiency of logistics claims processing and quality tracking. Learn more about our B2B lithium-ion battery ocean freight safety monitoring solutions now!

Read more
Hongke Dry Goods

[Hongke Solutions] Cybersecurity Training for Hong Kong’s Financial Industry: From Compliance Records to Human-Factor Risk Management

Cybersecurity Training Solutions for Hong Kong Financial Institutions (Banks, Securities Firms, and Insurance Companies). Designed to meet the compliance requirements of the Hong Kong Monetary Authority (HKMA) and the Securities and Futures Commission (SFC), these solutions help prevent Business Email Compromise (BEC), OTP phishing, and deepfake attacks. Utilizing KnowBe4’s automated simulation exercises, this solution establishes a human-factor risk management mechanism with a complete audit trail.

Read more
Hongke Case

[Hongke Applications] Hongke IDS Multi-Camera Imaging Solution – Overcoming the Challenge of Precise 3D Observation of High-Speed Dynamic Plasma

Hongke, in collaboration with IDS, offers a multi-camera synchronized imaging solution featuring microsecond-level ultra-short exposure and back-illuminated global shutter CMOS sensors, successfully overcoming bottlenecks in high-speed, microscale plasma 3D point cloud reconstruction and dynamic observation. Learn more about Hongke’s industrial machine vision solutions today!

Read more

Contact Hongke to help you solve your problems.

Let's have a chat