Spatial Intelligence is currently one of the most prominent technologies, and at its core is location recognition technology. Visual Localization (VL) is an AI technology that uses visual data to determine locations by utilizing cameras on robots or smartphones. Unlike GPS, VL works accurately indoors without the need for expensive equipment like laser scanners, making it a versatile solution for robots or AR. NAVER LABS has been focusing on VL research for several years and applying the technology in a variety of fields.
This article provides an in-depth look into the key features and applications of NAVER LABS’ VL technology.
Fast and Accurate Location Detection with Just One Photo
VL works as follows: First, it creates a digital twin—an exact replica of a space. Then, the system extracts key features from the 3D spatial data: edges, points, and patterns. When a photo is taken using a camera, AI compares these features with the photo to identify the current location.
Location detection is not very useful if it is slow or imprecise. However, in the demo below, you can watch the system instantly detect the location when a photo is taken with a tablet. The 6-DoF (Degree of Freedom) position, including orientation and posture, is displayed as a green square. The processing speed is within 0.3 seconds with a location error margin of less than 15cm and a rotation error margin of 3 degrees, making it highly accurate.
Beyond GPS: Indoor and Complex Environment Adaptation
As demonstrated in the demo above, one of the key strengths of VL is its ability to function indoors where GPS is unavailable. Its accuracy surpasses that of GPS, making it extremely useful for autonomous robots such as humanoids and quadrupeds, as well as for enabling devices like smartphones and AR glasses to accurately recognize current locations both indoors and outdoors.
Because VL relies on visual data, some may wonder whether it is vulnerable to changes in lighting, weather, or seasons, or whether its recognition ability decreases in crowded environments or when experiencing interior changes. This is where technological superiority comes into play. NAVER LABS has continuously advanced our AI models used in VL to ensure stable localization even through a variety of environmental changes. As a result, NAVER LABS’ VL works smoothly outdoors with moving objects as well as day-to-night light changes and geographical changes.
From Digital Twin to AI: Full-Stack VL
What makes NAVER LABS’ VL so exceptional? The biggest reason lies in our full-stack technology. NAVER LABS owns and controls every layer of the technology stack, including mapping devices for digital twins and algorithm development, AI model training, cloud processing environments, and the application of VL in various machines and services.
One significant advantage is that NAVER LABS can directly test and apply VL to our self-developed machines and services, including robots, autonomous vehicles, and AR, where precise positioning is critical. This eliminates the need to rely on external solutions or equipment and allows for faster development, application, and upgrading of our technology. As a result, NAVER LABS continues to be recognized as a global leader in VL technology in related industries as well as the academic world.
VL as a Cloud Service
We at NAVER LABS have launched our AI-based localization product, called ARC eye, through the NAVER Cloud Platform. With ARC eye, anyone can integrate VL into service robots, mobility solutions, and AR services—as long as a camera is provided.
This technology is crucial as it enables precise positioning indoors and outdoors, setting the foundation for a wide variety of future services. While VL is an inherently powerful technology and already one of the most advanced AI localization technologies available, we at NAVER LABS are enhancing our speed, performance, and usability at an accelerated pace each year. This AI system perceives space as accurately as the human eye, and it is a key technology that will be integrated into future platforms.