Gaode, create the next generation AI in the real world.

CN
1 hour ago

"Spatial Intelligence" is no longer just about more information itself, but rather experiences and judgments that are closer to the current situation and action arrangements.

Text Han Songyang

AI has spent years learning to fluently talk about the world. Now, it must face a far more difficult task: truly entering the world.

It is not hard for large models to create a weekend plan. They can recommend ten restaurants in seconds, summarize online reviews, and generate a seemingly organized tour route. However, when a person actually leaves their home, the questions quickly change: Are the recommended shops open today? Will there be a line if I go now? Where should I park? Which exit should I come out of the subway station? Is the entrance on the map under construction? And so on…

These questions are hard to answer solely with language; the real world is not a text that has already been written and is waiting for the model to summarize. Roads can become congested, businesses can relocate, foot traffic can ebb and flow, and weather and time can change the status of the same location; each choice made by a person will push this world into a different next moment.

This is also why, after language models, more and more AI researchers are turning their attention to "spatial intelligence."

Fei-Fei Li, a professor at Stanford University and co-founder of World Labs, states that spatial intelligence is a foundational architecture of human cognition. It relates not only to our ability to recognize objects, understand our environment, and judge spatial relationships such as distance and direction, but also supports our perception, reasoning, planning, and actions in the real world.

It is through this capability that humans can build cognitive models of the world in complex environments, predicting changes, solving problems, and interacting with their surroundings.

If large language models are breakthroughs in the first phase of artificial intelligence, then spatial intelligence may be a more important issue and a challenge that needs to be addressed next.

As AI Begins to Face a World That Cannot Be Generated

The training of large language models has a relatively clear goal: predicting the next word based on the preceding text. The vast amount of text available on the internet also provides ready-made training materials for this task.

However, the real world does not have such unified and clear data forms. Machines not only need to recognize objects; they must also understand the distances, directions, and occlusion relationships between them; they must not only see the immediate scene but also remember the spatial structure to judge how time and actions will change it.

As the scene expands from a room to a street to a city, the complexity increases. Buildings, roads, and places form a three-dimensional space, while traffic, foot traffic, weather, and business operations cause this space to change constantly. What machines need to understand is not just "what is here," but also "what is happening now" and "what might happen next."

Consequently, world models have become an important technological path to achieving spatial intelligence. They attempt to allow AI to internally form a model of the external world to represent space, simulate changes, and plan the next actions. Simply put, spatial intelligence is a capability that machines aspire to acquire, while world models are one of the methods to realize this capability.

Around this goal, different companies have chosen different entry points.

World Labs, founded by Fei-Fei Li, starts from the generation and reconstruction of three-dimensional spaces. The Marble, launched in 2025, can generate a three-dimensional world that users can enter, explore, and edit based on text, images, videos, or rough three-dimensional structures. On September 1, 2026, World Labs released the next-generation world model Atlas, which places text, images, videos, and three-dimensional information into the same spatial context, generating new observational perspectives, reconstructing realistic scenes, and simulating spatial changes over time.

Google DeepMind has chosen a different approach. Its Genie 3 can generate an environment in real time based on textual descriptions, allowing exploration and manipulation. As agents move within it, the model needs to continuously generate the next frame of the world, remember past scenes, and simulate how the environment will change based on the agent's actions.

NVIDIA's Cosmos serves more directly in "physical AI," such as robotics and autonomous driving. It can generate future scenarios under varying weather, lighting, road conditions, and unexpected events, providing machines with training and testing environments. Situations that are rare, expensive, or even dangerous in reality can be repeatedly simulated in the world generated by the model.

Amap has taken a different technical paradigm.

If we consider all businesses related to three-dimensional generation, Robotaxi, and embodied intelligence, Amap's technology system also involves the expression, simulation, and action planning of the world. However, Amap emphasizes that spatial intelligence is not about generating a world from scratch, but understanding an already existing, continuously changing world where numerous real actions occur daily.

This relates to Amap's past accumulation. Unlike other artificial intelligence laboratories, Amap, as a consumer product with over a billion users, has a long-term record of the real world, such as how roads connect, where buildings and businesses are located, how traffic changes, what destinations people choose, and which routes they travel.

What it aims to solve is not what a place can be generated into, but rather what its current state is, what might happen next, and how users should act.

This also brings forth different evaluation standards. For generative world models, a space that has not been photographed can have multiple reasonable answers; however, in maps and navigation, there is usually only one correct answer in reality.

Why Should Maps Stand at the Entry Points?

In the past, maps mainly recorded places and roads, while navigation calculated routes from point A to point B. Amap CEO Guo Ning summarizes these two capabilities as "connecting the real world." In the AI era, he hopes Amap can further achieve "understanding the real world."

He stated that connection solves "how the world arrives," while understanding addresses "what the world is truly like and how it changes."

Guo Ning breaks down the ability of a map company to understand the real world into four layers.

The first layer is the connection relationships of the two-dimensional world. Roads, subways, pedestrian networks, mall corridors, parking lots, and building entrances together form a vast network. A location is not just the latitude and longitude on the map, but its relation to the surrounding space: which street it connects, which subway entrance is the closest, which direction is easier to enter, and where to go after eating.

The second layer is three-dimensional space. The height of buildings, width of roads, floor relationships, occlusions, and views allow machines to further know what the real world "looks like." Maps change from flat points and lines to a three-dimensional space that can be accessed and observed from different positions.

The third layer is time. The real world is not static. The same restaurant may be in completely different states on a weekday afternoon and a weekend evening; the same road will show varying traffic efficiency during rush hours, rainy days, and holidays. Incorporating time transforms the map into a continuously changing spatiotemporal system instead of a static space.

The fourth layer is the "flow" in the real world. People move, cars move, and the active areas of the city continuously shift. Information such as foot traffic, car traffic, searches, navigation, visits, and stays enables Amap not only to observe the space itself but also to see how people interact with that space.

These four layers form a progressively expanding framework: the two-dimensional network shows how spaces connect, three-dimensional data supplements the physical structure of the world, time allows static spaces to start to change, and foot traffic, car traffic, and user actions present how this world truly operates.

"When the two-dimensional network, three-dimensional space, temporal changes, foot traffic, car traffic, and navigation behaviors are unified, what we obtain is no longer a map, but a city world model capable of prediction and simulation," Guo Ning said.

This analogy redefines the boundaries of a map company's capabilities: it should not only be a database that stores roads and places, but can also become a continuously updated "model," an entry point for understanding how the world operates.

XU Mu, head of Amap's spatial intelligence technology, borrowed Fei-Fei Li's categorization of world models—Renderer, Simulator, and Planner—to explain the division of labor and collaboration among different capabilities in Amap's spatial intelligence system: from presenting the world, to extrapolating changes, and finally guiding actions.

The Renderer is responsible for expressing or generating an observable three-dimensional world with high precision. Amap's recently landed ABot-World and ABot-Earth series models embody this capability and support the three-dimensional representation of maps and aerial street scenes.

The Simulator constructs computable and interactive environmental states, such as simulating how congestion spreads in surrounding road networks after a road is temporarily closed, or how different diversion measures affect foot traffic distribution in commercial areas during holidays.

The Planner combines environmental observations, task goals, and available state predictions and simulation results to generate action plans under environmental and task constraints. This can manifest as travel route recommendations for ordinary people, or extend to autonomous navigation and physical operations of robots.

Supporting the collaborative operation of this system is a MoT (Mixture of Transformers) architecture capable of collaboratively processing diverse and heterogeneous information. Its core balances specialized processing within modalities and information fusion between modalities, converting multi-source data into comprehensive representations for different tasks.

Below this lies Amap's long-term accumulation of extensive real-time spatiotemporal data—this data connects physical space, dynamic changes, and real actions, providing continuous and rich evidence for models to understand how the world operates.

According to Xu Mu's classification, Amap's data comprises four types. First is spatial geometric data, which describes the morphology and structure of the physical environment such as terrain, buildings, and roads; second is spatial semantic data, which describes the traffic attributes of roads, the functions of places, and the service information of businesses; third is dynamic status data, which describes foot traffic, car traffic, traffic conditions, and environmental changes over time; and fourth is behavioral interaction data, which correlates a person's actions, movement processes, and feedback results in specific environments. This information collectively forms a multidimensional data foundation, ranging from spatial structures to environmental changes and action outcomes.

Beyond data, Amap's other layer of barrier is its long-term established data and AI infrastructure.

This means that when new technical paradigms emerge, Amap can rely on its mature data collection and production systems to quickly acquire, process, and organize massive high-precision data that adapts to new methods, and accelerate the development and application verification of spatial intelligence and physical intelligence models through large-scale training, evaluation, and deployment capabilities.

More critically, there is a feedback loop formed between applications, models, and data.

Amap's models do not search for use cases only after completing training. Services such as maps, navigation, traffic, local living, and digital twins continuously generate new demands, data, and feedback. When new technologies enter real business scenarios, user and customer interactions generate new feedback, prompting the next iteration.

For instance, in robot navigation, the system can identify capability gaps from failure cases in complex environments, utilizing scene generation and simulation to create targeted data, and through training, independent evaluation, and real-world validation, continuously improve the model's generalization ability. The improved model can then enhance scene generation and evaluation capabilities, creating better conditions for the next round of iterations.

This "capacity enhancement drives the improvement process upgrade" mechanism also provides a foundation for exploring RSI (Recursive Self-Improvement), allowing spatial generation, environmental simulation, and action planning to mutually promote and continuously expand the boundaries of spatial intelligence.

When "Spatial Intelligence" Enters Real Life

The evolution of any technological foundation ultimately needs to answer how it changes the everyday experiences of ordinary people.

In the past, maps typically began working after users determined their destination: users first decided where to go, then opened the map to search for a route. Amap now wants to take a step forward and backward—helping users choose and judge before departure, adjusting actions based on real-time changes on the road, and continuing to understand the surrounding space upon arrival.

Building on the aforementioned real-world model, Guo Ning believes that Amap also needs to calculate four things that are closer to user experiences: trust, taste, presence, and context.

Trust answers which places, evaluations, and recommendations are worth believing; taste responds to whether a place is suitable for specific individuals; presence allows users to understand what experiences they will have when immersed in that environment; and context assesses what the appropriate choice is in this specific time, place, and trip.

A complete offline action by a user can illustrate how Amap uses products to accomplish this calculation.

The Street Ranking is the starting point of this chain. Its initial question is not "how to go," but "where to go," trying to solve how to gain trust in a list.

Traditional rankings primarily rely on online information such as ratings, reviews, and clicks. Amap, however, aims to judge whether a shop is truly worth visiting based on users’ actions in the real world. Yet, a visit itself carries different meanings: an individual may traverse half a city just to dine at a shop, or simply pass by on the way to work; they might have visited just once or live nearby and be a repeat customer.

In 2026, the Street Ranking further incorporated these differences into its algorithms. It differentiates between dedicated visits and walk-bys with the "proportion of dedicated visits," while measuring repurchase and local user long-term choices with "repeat customer proportion" and "local resident proportion." When real visits and long-term choices can mutually validate each other, a shop's recommendation becomes more credible.

Of course, a real visit doesn’t inherently equate to genuine reputation. Merchant locations, traffic convenience, and operating hours can all influence visit behavior. What the Street Ranking does is add more reality signals beyond pure traffic, combining actions and evaluations to get as close as possible to the real weight of a choice.

Building upon trust, the Street Ranking also seeks to further understand "taste."

Whether a shop is worth visiting is not the same answer for everyone. Hence, Amap's Street Ranking 2026 introduces professional ratings. For more specialized rankings, such as those for cafes, the input from industry experts will carry more weight in the scoring—if two cafes exist, one recognized by more coffee aficionados should obviously receive a higher score than a mainstream chain brand.

After deciding where to go, the next question users face is what that place is actually like. Aerial street views fulfill the spatial preview before departure, corresponding to Guo Ning's "presence."

The first-generation aerial street view can overlook neighborhoods from the air along a preset path, identifying shop facades with some scenes allowing indoor exploration. Aerial Street View 2.0 further transforms fixed viewing into free exploration: users can adjust position, direction, height, and observation angles, getting a sense of the destination and its surrounding environment in advance.

Its significance lies not only in making maps more realistic. For an unfamiliar neighborhood, tourist attractions, or complex building entrances, two-dimensional maps can only inform users where the destination is, while three-dimensional space allows people to imagine "what it would feel like to be there": what roads are around, what kind of environment the shops are in, how to enter, and what can be seen.

The Avoidance Guide mainly addresses "context," dynamically verifying a plan by integrating operating hours, travel pace, estimated arrival time, weather, costs, how convenient the route is, and parking information.

A shop may be worth visiting and align with user taste, but if it is closed upon arrival, or if the journey takes significantly longer than expected, this recommendation still fails to complete its task. Understanding context entails placing a location back into a specific time, place, and itinerary, and reassessing whether it still represents an appropriate choice.

Once into the travel itinerary, Navigation Live takes over.

Amap Navigation Product Lead Sun Chong summarizes the changes in navigation into four evolutions: First, navigation evolved from paper maps to digital algorithms, learning static pathfinding; second, from static road networks to real-time foot traffic, learning to avoid congestion; the third evolution upgraded from two-dimensional planes to three-dimensional twins, learning to clarify the world; and now, it has moved from guiding a segment of the journey to taking charge of an entire day, learning to accompany life.

This shift is first reflected in the questions users pose. A request like "Help me find a Huaiyang restaurant along the way, within an average of 200 yuan, that is still open when I get there," incorporates conditions of location, price, route, and time. Traditional navigation typically requires users to first search for restaurants, check information for each, and then return to the map for route comparisons; Navigation Live attempts to retain these conditions continuously, directly completing the selection and planning.

It relies on an ever-updating "dynamic spatiotemporal context": where users are at that moment, which route they are taking, the state of surrounding roads and locations, what tasks have been communicated in advance, and what information is being input through cameras and voice.

At this point, the four elements mentioned by Guo Ning also begin to coalesce into the same product. Navigation Live needs to utilize trusted location and road information, understand the preferences previously expressed by the user, establish situational awareness through the camera, and ultimately judge based on the context of that time, place, and entire trip.

This is also the key difference between Navigation Live and general visual question-answering. A general visual model can recognize a building or intersection in a scene but may not be able to identify which location they correspond to on a real map. Navigation Live must correlate the buildings, roads, and entrances captured by the camera with users' locations, directions, Amap's road network, and place information.

Through this correspondence, "that building" in the scene becomes a specific shopping mall on the map, and "the small door next to it" will carry the attribute of whether it can be accessed. The real physical spaces users see through their mobile cameras thus become searchable and navigable objects.

From the Street Ranking, Aerial Street View, and Avoidance Guide, to Navigation Live, Amap aims to connect an offline action: trust helps users judge what is worth visiting, taste makes recommendations more aligned with specific individuals, presence allows people to understand a space in advance, and context informs how to act at that moment.

In other words, spatial intelligence provides users with a more personal perception, no longer just about more information itself, but rather experiences and action arrangements that are closer to the current situation.

Spatial Intelligent Agents, Amap’s Imagination

Transforming demands from a linguistic world into a set of actions that can be executed in the physical world, the Street Ranking, Aerial Street View, Avoidance Guide, and Navigation Live represent Amap's current product answers in spatial intelligence, but it is evidently not the endpoint of spatial intelligence.

If this capability continues to evolve, in the future, users may no longer need to switch back and forth between rankings, street views, routes, and guides. The system will continuously comprehend the intentions of the same individual, along with their context in time, location, and environment, and invoke different capabilities based on real-time changes.

At that time, what users encounter might not be several independent functions but an intelligent agent that always knows "where you are, what you are preparing to do, and what you might encounter next."

Guo Ning prefers to understand Amap's changes as a "paradigm shift in capabilities" rather than just a technical upgrade in maps and navigation.

In the past, Amap primarily optimized a specific location or route, whereas spatial intelligence must determine what choices are more appropriate under concrete individuals, times, locations, weather, and preferences. This means the value of a location is no longer a fixed score. The system must not only know "what places are good," but also judge, "is it suitable for this person now?"

Looking further, a trip or weekend excursion is not merely ten independent locations. Where to stay, what time to leave, which path to take, where to pause, when to eat, and where to go afterward all influence each other. The previous choice alters the time, location, and state of the subsequent choice.

Guo Ning believes Amap ultimately needs to consider these locations and actions as a continuous sequence. The AI should not just provide a list of destinations or the fastest route but should offer a set of action suggestions that more closely align with users' specific needs while continuously adapting to real-world changes.

Sun Chong, the navigation product leader at Amap, summarizes this shift as moving from "issuing a series of short commands" to "delivering a long-term intention."

When users say, "keep an eye on my itinerary for the next two hours," the AI must recall prior requests, understand the user's current location and status, and appropriately manage changes that occur along the route at the right times. Users will no longer need to break tasks down into multiple steps like searching for locations, checking operating hours, comparing routes, and finding parking.

As part of this process, Amap's product role also begins to shift. It no longer simply awaits user input, but rather resembles an active and continuously operating agent in the real world.

This expands Amap's imaginative space beyond the map app—to become the next generation of AI in the real world. Location recommendations, spatial understanding, and action planning can appear on different terminals like vehicles, smartwatches, smart glasses, and robots. The form of devices can change, but the underlying problems to be solved remain the same—executing intended actions from the digital world in the real physical world.

However, as maps begin to inform users' choices, the complexity of the issues will also increase. The deeper the system intervenes in real actions, the higher the demands for accuracy and credibility. Amap must not only enhance its system's capabilities but also enable it to discern when to make judgments and when to exercise restraint.

Real-world data, mapping systems, and high-frequency application entry points provide it with a unique starting point, but whether these advantages can be organized into a stable, credible, and coherent user experience still needs to be tested through repeated real actions.

Cover image source: "Minority Report"

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink