Because of this feature, Gaode has become prominent.

CN
1 hour ago

AuthorSun Rui
WeChatlombredeau

AI can now arrange a trip clearly: where to go, what to eat, and where to stay, all laid out in detailed guides.

However, when entering a shopping mall and looking for a restroom just a few dozen meters away, you may still need to confirm back and forth: which passage to take, whether to go upstairs, and where the entrance really is.

For traditional navigation, it doesn't matter whether the phone is horizontal or vertical, since the main tasks are three: reporting intersections, calculating distances, and advising to turn around when you go wrong.

But Navigation Live has revolutionized traditional navigation. It first 'raises up' the phone—taking the phone off the stand and holding it vertically, the camera starts to observe the road, and combined with location and direction, translates the location of the restroom into visual guides that correspond to the current environment, thus aligning the digital information on the screen with the real world in front of you.

On September 10, Amap released the Street Scanning Rank 2026 and introduced the Navigation Live function. Sun Chong, head of Amap's navigation product, stated in an interview with Silicon Star: “The essence of Navigation Live is an embodied intelligent agent equipped with ‘spatiotemporal context + on-site perception + end-to-end action capability’. It turns the user's phone into a search box for the real world, capable of recognizing and interpreting the buildings, attractions, shops, and environmental information in front of them.”

It sounds like navigation has been fitted with "AI eyes," but "seeing" is just the first step.

For instance, in a historic downtown of an unfamiliar city, driving or walking to three adjacent narrow alley entrances, traditional navigation might only remind you to “turn right in 50 meters.” The direction may be correct, but which entrance it is often still requires guessing.

What Navigation Live aims to do is to turn a right turn on the map into a recognizable path in front of you: enter from the right side of the convenience store with red lanterns, paying attention to the delivery van unloading on the left. If you change your mind and want to try a different local cuisine, it can continue the previous conversation to find a restaurant and replan the route according to your current location and road conditions.

Holding the phone upright is not difficult. The challenge lies in making it aware of what it sees, where you are, and what to do next; Amap's goal is to enable navigation to understand the real world.

1

How Amap enables navigation to understand the real world

Navigation Live may seem like just adding a camera to navigation, but the real difficulty is not in having AI recognize an image, but in understanding the complex dynamic spatial relationships behind the image.

It needs to simultaneously understand where the user is, which way they're facing, what is in front of them, and the relationships between these objects and the roads, buildings, and places.

For example, recognizing that there is a restaurant in front of you is just visual recognition; further confirming which POI on the map the restaurant corresponds to, how far away it is from the user, and where the entrance is, involves comprehensive judgment of location, mapping, and spatial relationships.

Behind these capabilities lies Amap's spatial intelligence at work. It can be simply understood in three levels:

The first is dynamic perception, combining camera images, location and sensor information, as well as real-time traffic data to help AI understand what is happening in front of the user;

The second is spatiotemporal reasoning, not just judging what is currently seen, but also predicting what might happen next by combining road and traffic conditions;

Finally, three-dimensional spatial representation accurately relates digital information about places, roads, and environments to real space, allowing AI's understanding to continuously maintain the correct position as the user moves.

Therefore, behind Navigation Live is a set of spatial intelligence capabilities connected in the sequence of “perception - understanding - reasoning - action.” The camera allows AI to see the real world, while location, mapping, traffic, and three-dimensional spatial data enable it to truly understand the position of what it sees, its relationship with the user, and what can be done next.

This also makes Navigation Live more like an intelligent agent that can continue to move along with tasks, rather than a simple question-and-answer visual entry point. It doesn’t just see the sign in the frame; it also considers where you are at the moment, which direction you are facing, whether you are walking or stopped, and what changes are occurring around the road conditions.

Thus, it doesn’t need to wait to be questioned. If a lane is about to narrow, there is a blind corner ahead, or you pass by a well-reviewed old shop, it can proactively provide remarks at the right moment instead of making one repeatedly issue commands like using a tool.

After recognition, it can continue to take action. General visual models identify an object and usually provide a description; while Navigation Live, after recognizing a location, can directly connect to Amap’s existing capabilities—recalculating the route, pointing to a specific entrance, checking further information about that shop, turning "recognition" into "what to do next."

1

Testing Navigation Live: Finding a restroom is no longer difficult

The above changes manifest more at the product logic level, so using Navigation Live in practice allows for a more intuitive observation of this change.

We conducted real-world tests in two common scenarios: finding a restroom indoors and selecting a restaurant.

Everyone often encounters the problem of finding a restroom indoors, which is a very frequent navigation need.

In typical cases, we need to open the map search and choose from the POIs (points of interest) provided by the navigation software. However, in complex spaces like malls and scenic areas, the position on the map does not necessarily translate directly into actionable instructions in reality.

The user knows the restroom is only a few dozen meters away but may still need to repeatedly judge between multiple entrances, stairs, elevators, and passages.

Navigation Live simplifies this step. By tapping on the camera icon next to the search box, AI captures the current environment and, combined with position and direction, translates the restroom POI on the map into visual guidance corresponding to the environment in front of them.

In this process, Navigation Live not only provides destination information but also attempts to establish a correlation between the map coordinates and the real space. For relatively complex scenarios like malls and districts, Navigation Live can help reduce search time and unnecessary detours.

The second scenario is a more complex yet equally frequent need when shopping: choosing a restaurant. When you feel hungry and tired during travel and want to find a restaurant nearby, multiple dining options on the street can make it difficult to choose at times.

In the past, you usually had to confirm the names of various shops first, then enter a map or review product to search for ratings, reviews, and recommended dishes.

Navigation Live offers a more straightforward interaction method.

We can point the camera at the street and ask Amap through Navigation Live “What is the rating of this shop?” or “Which shop is better for a solo meal?”

Navigation Live can directly recognize the restaurant within the frame, and clicking on the street scanning rank icon will provide more restaurant ratings, average consumption information, and more.

This effectively correlates visual, location, and consumption information all at once: first recognizing the restaurant in front of you, then associating it with the specific POI on the map, and retrieving evaluation information, organizing answers according to your questions. The original lengthy process of “recognizing objects - searching objects - obtaining information” is compressed into a single query.

Finding a restroom resolves the relationship between the current location and the next action, while asking about a restaurant resolves the relationship between real objects and digital information; both together answer the same question: How does AI enter and understand the real scene you are experiencing.

1

Spatial intelligence reshaping a whole day's life experience

Of course, Amap's spatial intelligence capabilities do not solely serve Navigation Live. The Flight Street View 2.0 turns destinations into a three-dimensional space that can be freely explored, allowing users to adjust viewing positions, angles, and heights, gaining an immersive presence before departure. The Street Scanning Rank combines signals from real visits, special trips, and repeat visits, along with authorized evaluation information to help users judge whether a place is worth visiting; the Avoiding Pitfalls Guide verifies crafted plans within specific time, space, and traffic conditions, checking whether the originally static itinerary is genuinely executable.

Behind these products lies Amap's gradual application of spatial intelligence across different segments of travel.

The ABot-Earth 0.7 showcased at the launch event is an intuitive representation of Amap's three-dimensional spatial expression capabilities. As a 3D native city world model, it organizes space, time, and real information into a continuously accessible, freely explorative, and interactive three-dimensional world.

From Street Scanning Rank to Flight Street View 2.0, Avoiding Pitfalls Guide, Navigation Live, and ABot-Earth 0.7, Amap is exploring the same question: how to align AI with the real world.

Today, generating an answer is no longer difficult; the challenge lies in ensuring the answer fits the time and place. Traveling has never required an answer detached from the scene but rather information relevant to this street, this entrance, and the current road conditions.

To achieve this, one end involves long-term accumulation and continuous updating of maps, roads, locations, and street views, while the other end transforms these data into spatial understanding capabilities. The former determines whether AI can access the real world, while the latter decides whether it can comprehend it.

The past maps answered the question of where the world is; now they need to address the relationship between this world in front of you and your next step. This is also why Amap serves as the gateway for AI into the real world.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink