ABot-Earth 0.7 – The World's First 3D Native Urban World Model Launched by AutoNavi

Executive Summary:
ABot-Earth 0.7 is the world's first 3D native urban world model introduced by AutoNavi Maps. Its core innovation lies in the use of a 3D native technology architecture, incorporating spatiotemporal da...
1. What is ABot-Earth 0.7
ABot-Earth 0.7 is the world's first 3D native urban world model introduced by AutoNavi Maps. Its core innovation lies in the use of a 3D native technology architecture, incorporating spatiotemporal data that includes both spatial and temporal information, enabling the model to develop a native understanding of three-dimensional space. This model supports users inputting a satellite image or a text description, and can generate a kilometer-level 3D city on a consumer-grade GPU within just 10 minutes, achieving an efficiency improvement of approximately 1000 times compared to traditional methods. The model covers over 196 countries and regions, supporting continuous generation across all scales—from planetary views to street-level scenes. Users can freely explore and interact in real-time, and it has already been deployed as the entry point for AutoNavi's "spatial intelligence" in the Flight Street View 2.0 feature.

Image source: Official article
Technical Positioning and Domain: Belongs to the domain of spatial intelligence and 3D generation, focusing on the real-time generation and interaction of city-level 3D scenes. This model integrates generative AI, 3D Gaussian Splatting (3DGS) rendering, and geographic information systems, positioning itself as a generative alternative for digital Earth, differing from traditional image reconstruction approaches that rely on stitching real-world footage.
Development Background: Developed by the AutoNavi Maps R&D team, this model is built upon a decade of spatiotemporal data accumulation, with AutoNavi's BeiDou positioning system being called nearly 10 trillion times daily, forming a unique and hard-to-replicate training dataset of the real world. The motivation for development stems from the limitations of traditional 3D city reconstruction, which relies on manual modeling and specialized data collection equipment, resulting in high costs, long development cycles, and limited coverage. AutoNavi aims to use generative AI to achieve large-scale, low-cost construction of urban spaces.
Core Value: Addresses three major pain points in traditional 3D city construction: high data collection costs, long update cycles, and lack of coverage in remote areas. Its end-to-end generation model changes the traditional "collect first, then reconstruct" workflow, enabling the instant generation of city-level 3D spaces at any location. Furthermore, it opens up engine integration capabilities for developers, establishing a complete path from user experience to commercial application.
Technical Features: Employs an end-to-end 3DGS generation strategy, outputting a 3D Gaussian Splatting format scene in one go, without the need for step-by-step reconstruction. The two-stage generation process first produces a sparse structure and then refines the geometry and attributes. The 3D native architecture establishes a fundamental understanding of three-dimensional space, differing from the post-processing method of elevating 2D images to 3D. It holds potential for maintaining spatial consistency and enabling dynamic perception evolution.
2. Key Features
Image/Text-to-3D City: Input a satellite image or a textual description, and the model generates a kilometers-scale 3D city scene within 10 minutes. This feature employs a two-stage generation process, first producing a sparse structure (occupancy layout), then refining it to generate geometry and attributes (active voxel), progressively enhancing details under the condition of the satellite image. It is suitable for quickly building a city prototype or validating creative scenarios.
Full-Scale Continuous Generation: A single model seamlessly generates scenes ranging from planets, cities, to street-level landmarks. As the viewpoint continuously zooms in, the spatial structure and visual representation remain highly consistent, eliminating the need to switch models or load different layers. This capability breaks through the disjointed feel caused by traditional multi-scale model stitching, delivering a truly continuous spatial experience.
Free Exploration Interaction: The model autonomously completes a full and continuous 3D space, allowing users to continuously enter, freely roam, and interact in real-time, without being restricted by the pre-defined routes and capture paths of traditional street views. This ability stems from the model's native understanding of 3D space, rather than the arrangement of existing captured materials.
Plug-and-Play Engine Integration: The generated results are output in 3DGS format, enabling seamless integration with Unity and Unreal Engine. Developers can directly add interactive logic, physics simulations, and particle effects within the game engine, making it applicable to game development, film pre-visualization, or digital twin production pipelines.
Flight Street View 2.0 Deployment: ABot-Earth 0.7 has been integrated as the underlying spatial generation capability into the "Flight Street View 2.0" feature of the Amap App. Users can control it with a joystick to explore complex buildings and large scenic areas with immersive 3D navigation, marking the model's transition from the lab to large-scale consumer applications.
Multi-Body Support: The official website's planet page supports browsing spatial data for Earth, the Moon, and Mars. Users can experience 3D navigation across planetary scales, providing a low-cost digital space foundation for scenarios such as astronomy education and preliminary deep-space exploration research.
Land Parcel Square Shared Ecosystem: The community section allows users to share and browse 3D land parcels generated by others, aggregating and re-consuming generated content. This makes it easier to visually assess generation quality and lays the groundwork for future creator ecosystems and content trading.
3. How to Use
Environment Requirements: ABot-Earth 0.7 provides an online experience mode, allowing users to access core features without requiring a local GPU or complex environment setup. Only a modern browser that supports WebGL is needed (recommended: Chrome 90+ or Edge 90+). Developers who wish to integrate the engine should prepare a Unity 2021+ or Unreal Engine 5+ environment to import 3DGS format assets.
Official Website Experience: Access the ABot-Earth 0.7 official website directly (link to be updated after official release).
Image-to-3D / Text-to-3D Operations: In the image-to-3D module, upload a satellite image (recommended resolution: 1024×1024 or higher), or switch to the text-to-3D module and input a description of a city (e.g., "a coastal mountainous city with high-density residential areas and a central business district"). After clicking "Generate," the system enters a two-stage generation process. After approximately 10 minutes, you can preview and explore the generated kilometer-scale 3D city in the preview window.
Experience within the AutoNavi App: Open the AutoNavi Maps app, search for scenic spots or complex structures that support Flight Street View 2.0 (e.g., large theme parks, ancient architectural complexes). Upon entry, the app will automatically load the 3D space generated by ABot-Earth 0.7. Use the virtual joystick to control the view direction and movement path, experiencing a continuous tour from an aerial perspective to internal navigation.
Engine Integration (for Developers): After generating a 3D scene on the official website, export the 3DGS format file and import it into Unity or Unreal Engine. Connect with the official rendering plugin (to be released), and add your own colliders, interactive logic, and lighting baking for use in game levels, film previsualization, or smart city visualization projects.
On a practical operation level, satellite images should ideally include clearly visible road grids and building layouts. For text descriptions, include key features such as terrain, building density, and functional zones, which can significantly improve the match between the generated scene and the intended outcome. The online experience is currently free, but individual generation tasks may require waiting in a queue. During peak hours, the estimated generation time may exceed 10 minutes.
4. Pros and Cons Analysis
| Pros |
|---|
| End-to-end 3DGS Generation Architecture: Outputs 3D Gaussian Splatting format scenes in one go, eliminating the need for step-by-step reconstruction or post-processing stitching. The rendering results maintain continuity and detailed expression, with quality significantly surpassing traditional voxel-based approaches. |
| Efficiency Leap by a Thousandfold: A consumer-grade GPU can generate kilometer-scale urban scenes in 10 minutes, improving efficiency by approximately 1,000 times compared to traditional reconstruction pipelines that require months, drastically reducing the time cost of spatial production. |
| Full-scale Spatial Consistency: A single model covers scales from planetary to street-level, maintaining structural continuity and visual consistency when zooming in, solving the scale discontinuity issues caused by multi-model stitching. |
| Global Coverage Leadership: Leverages AI generation to support expandable coverage across 196+ countries and regions, unrestricted by the scope of real-world data collection. It maintains urban generation capabilities even in remote areas. |
| Gaode Spatiotemporal Data Barrier: Decades of spatiotemporal data accumulation, with nearly 100 billion daily calls to the BeiDou positioning system, have created a unique real-world sample database that is difficult to replicate, offering a distinctive data moat for model training. |
5. Comparative Analysis with Similar Tools
| Comparison Dimension | ABot-Earth 0.7 (AutoNavi) | Google Earth / 3D Tiles (Google) |
|---|---|---|
| Core Architecture | 3D native generation architecture, trained on spatiotemporal data, end-to-end output of 3DGS scenes | Image reconstruction architecture, satellite imagery + aerial photography + point cloud stitching, relies on photogrammetry pipeline |
| Generation Efficiency | Generates kilometers-level 3D cities in 10 minutes on consumer-grade GPUs, efficiency increased by about 1000 times, requires task queue waiting | Requires professional acquisition equipment and months-long reconstruction pipeline processing, cannot generate instantly |
| Coverage Scope | Covers 196+ countries and regions, AI-generated models can be extended to any area, no blind spots in data collection | High-definition coverage of major global cities, low precision or missing models in remote areas |
| Interaction Method | Free entry, continuous roaming, real-time interaction, provides a sense of "presence" | Primarily browsing, with view angles and routes restricted by the data acquisition path |
| Openness | Generated results can be exported as 3DGS and integrated into UE/Unity, open to the developer ecosystem | 3D Tiles provides API, but scene generation capabilities are not open to third parties |
| Data Sources | AI-generated + autonomous completion, not limited by data collection materials, based on AutoNavi's BeiDou spatiotemporal data | Relies on real-world imagery updates, long update cycles and high costs |
Selection Recommendations: For developers who need to quickly generate 3D city scenes for any location and deeply integrate them into game engines, ABot-Earth 0.7 is currently the only option in the market with minute-level generation capabilities and open interfaces. It is especially suitable for game prototype validation, film and television virtual location scouting, and low-precision digital twin foundation construction. For professional GIS users requiring high-accuracy real-world geographic information globally, Google Earth remains the benchmark for data accuracy and coverage maturity. However, its limitation of not being able to generate new scenes makes it more suitable for browsing and annotation.
For cost-sensitive small and medium enterprises with scenarios focused on already-covered cities, ABot-Earth 0.7's free online generation mode is more affordable than the per-scene billing model of commercial remote sensing platforms. Note that the current version's generation accuracy is still insufficient to support high-precision mapping and key scenario validation for autonomous driving. Such professional needs still rely on traditional data acquisition and reconstruction methods or await future model updates.
6. Editor's Summary
The technological innovation of ABot-Earth 0.7 lies in introducing an end-to-end 3DGS generation strategy into city-level spatial construction, replacing the traditional photogrammetry reconstruction pipeline with a two-stage generation process, fundamentally changing the production paradigm of digital Earth. Its technical approach of "spatiotemporal data training + native 3D understanding" has a clear evolutionary direction. The efficiency metric of generating kilometer-level scenes in just 10 minutes on consumer-grade GPUs is competitive among similar solutions, and it has already undergone large-scale application validation in AutoNavi's Fly Street View 2.0.
In terms of practical value, this model reduces the production cost of 3D cities by several orders of magnitude, making high-precision urban modeling capabilities previously accessible only to large enterprises available to mid-sized developers and independent creators. This lowered barrier may foster a new spatial application ecosystem. Its coverage across 196+ countries and regions fills the gaps left by traditional digital Earth solutions in remote areas, providing a unified spatial foundation for global applications.
In terms of target users, game developers can regard it as a productivity tool for rapidly generating open-world city levels, autonomous driving companies can use it to build low-cost simulation training environments, and the cultural tourism and real estate industries can leverage it to quickly set up online 3D display spaces. AutoNavi's implementation in navigation scenarios gives it consumer-facing user experience value, going beyond mere technical demonstrations.
For future iterations, we look forward to the official team continuing to invest in improvements in generation resolution, support for dynamic elements, and the refinement of developer toolchains. These three directions will determine whether the model can truly evolve from a "city space generator" into a "city spatiotemporal simulation engine." Its long-term value depends on AutoNavi's ability to continuously convert its accumulated spatiotemporal data into model capability advantages and to create a developer network effect through open ecosystem construction.
7. Application Scenarios
Smart Mobility Navigation: Targeting users of AutoNavi maps and in-vehicle systems, ABot-Earth 0.7's generated 3D space can be directly embedded into navigation workflows, enabling 3D real-world navigation and destination pre-exploration. Users can virtually walk through scenic areas, commercial districts, and complex buildings before departure, reducing navigation costs caused by unfamiliar environments.
Autonomous Driving Simulation Training: Rapidly generate kilometers-level realistic urban scenes to provide a low-cost, customizable simulation testing environment for autonomous driving algorithms. This addresses industry challenges such as limited real-world testing scenarios and difficulty in reproducing extreme conditions, significantly enhancing the diversity of simulation training environments.
Gaming and Film Production: Leverage text-to-image and image-to-image generation capabilities to produce 3D city assets within minutes, which can be directly imported into Unreal Engine or Unity as base models for game levels or foundational elements for virtual film production. This greatly shortens the pre-production asset creation cycle for open-world games and virtual filming.
Digital Twin and Smart Cities: Construct a real-time interactive 3D city foundation to support urban planning scenario simulations, traffic flow modeling, and emergency evacuation drills, providing city managers with an intuitive, low-cost spatial decision-making aid.
Cultural Tourism and Real Estate Display: Cultural sites and museums can quickly generate online 3D roaming spaces for remote previews by tourists, while real estate projects can create immersive 3D model rooms for online property viewings. This reduces the cost of physical visits while expanding reach.
Embodied Intelligence Training Data Supply: Provide large-scale 3D scene training data for spatial perception algorithms in embodied robots and smart glasses, supplementing the diverse spatial samples that are difficult to obtain in real-world environments, thereby enhancing environmental understanding and navigation decision-making capabilities.
8. FAQ
Q: Is the city generated by ABot-Earth 0.7 completely consistent with real geographic information?
A: Not entirely consistent. The 3D cities generated in the current version are based on satellite images and textual descriptions, presenting a visual scene that aligns with the spatial layout characteristics of real environments, but it is not centimeter-level precise mapping. There are certain deviations in building height and road orientation from the real environment. It is suitable for browsing, demonstrations, and simulation training, but not appropriate as a basis for high-precision engineering measurements.
Q: What are the specific GPU requirements for generating KM-level cities in 10 minutes on consumer-grade GPUs?
A: The official has not disclosed the minimum configuration details, but the "consumer-grade GPU" typically refers to NVIDIA RTX 3060 and above. Actual generation time is affected by VRAM capacity, scene complexity, and server queue status. Detailed local inference configuration requirements are yet to be officially released. For now, it is recommended to prioritize using the official online experience.
Q: Can the generated 3D scenes be used directly in commercial projects?
A: In theory, the 3DGS format can be exported and integrated into UE/Unity for commercial use. However, the official has not yet released complete commercial licensing terms or technical documentation. It is advised to monitor updates to the official license agreement before commercial use, or contact the AutoNavi business team to confirm the scope of authorization.
Q: How does ABot-Earth 0.7 integrate with the navigation features of AutoNavi Maps?
A: Currently, the integration mainly occurs through the "Flight Street View 2.0" entry in the AutoNavi App, allowing users to explore complex buildings and large scenic areas in 3D. There is no official timeline for deep integration with real-time navigation routes, and future versions may embed 3D scenes into the navigation process.
Q: Does the current version support model fine-tuning or private deployment?
A: The official has not yet opened model weight downloads or private deployment solutions. Users can currently only generate content through the official online service and export it in 3DGS format for use. The ability to train or fine-tune the model itself is not available to the public.
9. Project Links
- Product Official Website: https://abot-earth.amap.com/
Related AI Model Articles

Ming-Image-0.1-Design: Ant Group Open-Sources 6B Parameter Image Generation Model, End-to-End Reimagining the Design Workflow
Ming-Image-0.1-Design is a 6B parameter image generation model open-sourced by Ant Group's InclusionAI team, specifically tailored for design scenarios. It supports 8K long, structured prompts and can...

Jev Chat Assistant – Open-Source AI Chat Companion for Generating the Most Appropriate Responses
Jev Chat Assistant is an open-source, non-intrusive AI chat assistance application that provides real-time reply suggestions in popular messaging scenarios such as WeChat, QQ, X, and Feishu. The tool ...
GPT-6 Luna: OpenAI's Cost Revolution and Capability Democratization in Lightweight Models
GPT-6 Luna is a lightweight AI model introduced by OpenAI, and it is a derivative version of GPT-6 Astra, alongside GPT-6 Sol. It is positioned for high-frequency, large-scale task scenarios. The mode...
Laya – Open-Source Non-Autoregressive AI Decision Model, an Open-Source Alternative to Jev
Laya is an open-source non-autoregressive AI decision model based on a bidirectional encoder architecture, developed by the ConvAI Innovations team. Unlike mainstream large language models, Laya does ...
© All Rights Reserved. Some content on this site is partially generated by AI with human review.
