Back to Model List

StepAmoo – A New Generation of Personal Agent from StepFun

AI Tech Editorial
RSS Feed

Executive Summary:

StepAmoo is a new generation of personal agent launched by StepFun, built upon the world's first agent-native operating system, Step AOS. It possesses operating system-level identity and system-level ...

1. What is StepAmoo

StepAmoo is a new generation of personal agent launched by StepFun, built upon the world's first agent-native operating system, Step AOS. It possesses operating system-level identity and system-level permissions. With five core capabilities – perception, memory, planning, connectivity, and execution – StepAmoo employs a dual-domain, three-step memory structure, enabling it to continuously understand user habits and preferences. Through an end-to-end cloud-edge协同 routing mechanism, StepAmoo can orchestrate atomicized services across applications to complete complex task loops, and seamlessly continue tasks across multiple devices. It aims to fundamentally redefine human-computer interaction from the ground up, becoming a truly 24/7 personal digital assistant.

Technical Positioning and Field: StepAmoo belongs to the agent technology domain, specifically positioned as an operating system-level personal agent. Unlike traditional voice assistants or AI applications, StepAmoo is not merely a functional entry point running on top of an operating system, but is natively embedded within Step AOS, granting it privileged access to directly schedule hardware resources, system services, and third-party applications. This enables it to execute complex, cross-application tasks, achieving a paradigm shift from "users seeking applications" to "applications seeking users," representing the evolution of personal AI assistants from tool-based to system-native intelligent agents.

Development Background: StepAmoo was developed by StepFun, a company founded by technical experts from top-tier tech firms such as Microsoft and Google, with deep expertise in foundational large models. The company has previously released a series of models, including Step-1 (a language model with hundreds of billions of parameters) and Step-2 (a multimodal model with trillions of parameters). The emergence of StepAmoo is closely tied to StepFun's agent smartphone, STEPX Neo, which is equipped with Step AOS and uses StepAmoo as its core interaction interface. This creates a complete closed-loop from foundational models, agent systems, to terminal hardware, showcasing the company's strategic layout in edge-side AI and operating system-level agents.

Core Value: StepAmoo addresses three major pain points of traditional voice assistants: limited permissions, lack of continuous memory, and the inability to complete complex cross-application tasks. By leveraging operating system-level permissions, it can directly invoke system atomic capabilities (such as location, camera, calendar, and third-party applications), combined with dual-domain long-term memory, to understand users' deeper preferences and automatically plan and execute multi-step tasks (such as automatically booking tickets, reserving hotels, arranging rides, and syncing schedules when planning a business trip tomorrow). This frees users from tedious manual operations, truly achieving "intent equals execution."

Technical Features: The technical architecture of StepAmoo comprises five pillars: dual-domain three-step memory (separating the user domain and agent domain, with recall latency as fast as 15 milliseconds on the memory-reasoning-memory chain), end-to-end cloud-edge协同 routing (dynamic switching between Fast Path, Deep Path, and Hybrid Path), cascaded model scheduling (edge-side Step Edge 300M-10B, cloud-side Step Flash 200B-400B and Step Pro 1.5T+), atomic capability engine (decomposing the system and applications into minimal atomic units, with orchestration made available through protocols like MCP and A2A), and four-dimensional security governance (trusted operations, visible processes, controllable permissions, and reversible behaviors). These technologies collectively support its efficient, secure, and cross-device agent capabilities.

2. Key Features

  • Multimodal Perception: StepAmoo supports three natural interaction modalities: voice, vision, and text. It can simultaneously understand voice commands, recognize image content (such as identifying objects or scenes through photo capture), and parse text input. This multimodal fusion perception capability enables the agent to accurately capture user intent even in complex environments (such as noisy backgrounds or multi-intent expressions), enhancing the naturalness and robustness of the interaction.

  • Dual-Domain Memory Accumulation: StepAmoo employs a unique dual-domain three-step memory structure. The user domain stores personal facts, situational preferences, and user profiles, while the agent domain accumulates domain-specific knowledge, cognitive patterns, and task methodologies. Through the "remember-analyze-recall" three-step process, it achieves automatic memory extraction, evolution, and efficient retrieval, with memory search latency as low as 15 milliseconds. This allows the agent to continuously adapt to user habits and provide increasingly personalized services.

  • Autonomous Task Planning: When a user expresses a complex intent (e.g., "Help me prepare for next week's birthday party"), StepAmoo utilizes a cascaded model scheduling approach. First, an on-device model quickly parses the intent, then a cloud-based large model performs in-depth task decomposition, autonomously generating an action plan with multiple sub-steps (e.g., searching for restaurants, sending invitations, ordering a cake, arranging transportation), and dynamically adjusting the plan to accommodate changes.

  • Atomic Service Connectivity: Through an atomic capability engine, StepAmoo decomposes the functions of the operating system and third-party applications into minimal atomic units (e.g., "send SMS," "create calendar event," "access camera," etc.). These units are then exposed via standard protocols such as MCP (Model Context Protocol), A2A (Agent-to-Agent), and CLI. Agents can freely compose these atomic services like building blocks, completing long-chain cross-application task loops without requiring user manual navigation.

  • Edge-Cloud Collaborative Execution: Based on task complexity, intent confidence level, network status, and privacy level, StepAmoo dynamically switches between three execution paths: Fast Path (edge-side closed-loop, suitable for simple real-time tasks, such as setting reminders), Deep Path (cloud-side inference, suitable for complex planning, such as generating travel itineraries), and Hybrid Path (edge-cloud handover, where the edge performs preprocessing, the cloud conducts deep processing, and the result is returned to the edge for execution). This mechanism achieves a balance between real-time performance and deep reasoning capabilities.

  • Result Interaction Delivery: Users only need to express their intent in natural language. After StepAmoo completes the full task chain, it proactively delivers the results in visual cards, voice announcements, or text summaries. Users can view the agent's complete execution steps and modify or revoke intermediate results, achieving a transparent and controllable interaction experience.

  • Four-Dimensional Security Governance: A security framework at the agent level is established, covering operational trustworthiness (all actions are based on user authorization), process visibility (users can view execution steps in real time), controllable permissions (permissions are granted on an as-needed basis and revoked after use), and reversible behavior (support for task revocation and result rollback). Full operation logs are auditable, and users retain complete control, effectively balancing agent capabilities with privacy and security.

3. How to Use

As StepAmoo is currently in the early co-creation and internal testing phase and is not yet available to the general public, the following guide outlines the expected usage based on the officially published architecture information and internal testing process.

  1. Obtain Internal Testing Access: Users must apply for internal testing access by following recruitment announcements posted on official channels of Jieyue Star (such as the official website, social media, or community). After submitting an application, users will need to wait for approval. Internal testing is typically open to developers, geek users, or specific industry partners, with limited spots available.

  2. Device and System Requirements: StepAmoo natively runs on terminal devices equipped with Step AOS. Currently known devices include the STEPX Neo, a smart agent phone launched by Jieyue Star. This device is specifically optimized for Step AOS and StepAmoo, featuring necessary hardware support (such as on-device AI chips and multi-modal sensors). Existing Android or iOS devices cannot run StepAmoo directly.

  3. Activation and Initial Setup: After logging into a Jieyue Star account on the STEPX Neo device, the system will guide the user through activating the StepAmoo agent. First-time users must grant permissions (location, photo gallery, calendar, address book, third-party apps, etc.) and set personal preferences (such as language, frequently used addresses, and interest tags). This information will be stored in the user's domain memory for future personalized services.

  4. Daily Interaction Methods: Users can activate the agent using a voice wake-up word (e.g., "Hi Amoo") or by long-pressing the device's side button. Once activated, users can describe their intent in natural language. For example: "Book a flight ticket from Beijing to Shanghai tomorrow afternoon, and choose a window seat." StepAmoo will automatically detect and initiate task planning, potentially confirming key information via voice or on-screen prompts during the process.

  5. Task Execution and Result Delivery: The agent schedules atomic services according to the plan and executes tasks across applications. Users can view the execution progress bar or step list in real-time on the device screen. After the task is completed, the results are presented in the form of notification center cards or full-screen summaries. Users can adjust the results via voice or touch (e.g., changing the time, canceling an order), and the agent will re-execute the task accordingly.

  6. Privacy and Security Operations: Users can view StepAmoo's operation history and manage authorized permissions in the settings, revoking access to specific apps or data at any time. For sensitive operations (such as payments or file deletion), the agent will require users to confirm twice, ensuring that actions are reversible.

Notes: Functionality and stability may be limited during the internal testing phase, and users are encouraged to provide timely feedback. Simple tasks processed by the on-device model do not require an internet connection, but complex tasks need a stable network to invoke cloud-based large models. It is recommended that users start with simple tasks during initial use and gradually allow the agent to learn their habits.

4. Pros and Cons Analysis

Pros
OS-level Identity: Natively embedded in Step AOS, with system-level permissions, enabling direct scheduling of hardware resources, system services, and third-party applications, achieving deep cross-application collaboration that traditional voice assistants cannot match.
Dual-domain Three-step Memory: Separates the user domain and agent domain, with a "remember-analyze-retain" chain enabling long-term memory accumulation, fast recall speed (15 milliseconds), and continuous adaptation to user habits, providing personalized services.
Edge-Cloud Collaborative Routing: Dynamically switches between edge-side fast response and cloud-side deep reasoning, balancing real-time performance and complex task processing capabilities, adapting to different network and privacy scenarios, and improving task execution efficiency.
Atomic Capability Scheduling: Decomposes the system and applications into minimal atomic units, allowing flexible orchestration and invocation, enabling long-chain task closure (such as travel planning, office workflows), far surpassing traditional app switching.
Four-dimensional Security Governance: Trusted operations, visible processes, controllable permissions, and reversible behaviors, providing agent-level security assurance. Users can audit all operations, enhancing their sense of control.

5. Comparative Analysis with Similar Tools

Dimension StepAmoo Traditional Voice Assistants (Siri/Xiaoai Voice Assistant) Huawei Xiaoyi (HarmonyOS)
System Identity Operating system-level native agent with system-level permissions, capable of scheduling hardware and system services. Pre-installed system application or functional entry point, with limited permissions and unable to directly invoke underlying system capabilities. Has certain system-level permissions within HarmonyOS (such as control center, device collaboration), but is not built on a native agent architecture.
Memory Capability Dual-domain three-step memory, long-term accumulation of user preferences and experiences, with recall as fast as 15 milliseconds, supporting continuous learning. Short-term command memory, no continuous context understanding, and each interaction is largely independent. Supports a certain degree of memory (such as schedules, preferences), but mainly single-domain, lacking knowledge accumulation across agent domains.
Cross-Application Scheduling Atomic service orchestration, capable of completing long-chain closed-loop tasks (such as travel, office work), and freely invoking atomic capabilities. Limited to calling specific apps (such as playing music, setting alarms), and difficult to coordinate across apps to complete multi-step tasks. Can collaborate across devices within the HarmonyOS ecosystem (such as between a phone and a tablet), but has limited atomic scheduling capabilities across applications.
Edge-Cloud Collaboration Dynamic routing relay between edge-side fast response and cloud-side deep reasoning, supporting three paths: Fast, Deep, and Hybrid. Primarily relies on cloud processing, with weak edge-side capabilities, and limited functionality when offline. Edge-cloud collaboration has certain capabilities (such as localized speech recognition), but task planning still mainly relies on the cloud.
Security Mechanism Four-dimensional security system: trusted operations, visible processes, controllable permissions, reversible behavior, with full auditability throughout. Basic permission management, lacks an agent-level governance framework, and users cannot finely control operation history. HarmonyOS security system provides permission levels, but lacks agent-level reversible behavior and audit capabilities.
Multimodal Perception Native support for multimodal fusion of voice, vision, and text, enabling natural interaction. Primarily voice-based, with partial text support, and limited visual capabilities (e.g., Siri can recognize images but not natively). Supports voice and vision (e.g., screen recognition), but the depth of multimodal fusion is less than that of StepAmoo.

Selection Recommendations: For users seeking an ultra-personalized intelligent assistant and willing to engage with the Step AOS ecosystem, StepAmoo is currently the most advanced choice in terms of technical architecture, especially suitable for high-frequency scenarios requiring complex cross-application task closure (such as business travel and office automation). For general consumers, traditional voice assistants (such as Siri and Xiaoai Voice Assistant) still meet daily needs in basic command execution, with mature ecosystems and widespread device adoption. Huawei users may prefer Xiaoyi, which has unique advantages in HarmonyOS multi-device collaboration, but its cross-application atomic scheduling and long-term memory capabilities are not as strong as StepAmoo. Google Assistant has broad integration within overseas ecosystems, but its system-level permissions and depth of edge-cloud collaboration fall short of StepAmoo. Overall, StepAmoo represents the evolutionary direction of personal agents, but its ecosystem maturity and device adoption still require time to develop.

6. Editor's Summary

StepAmoo demonstrates significant innovation in technological advancement: it elevates agents from the application layer to the operating system layer, achieving true native integration at the system level through Step AOS. This architectural design does not simply overlay AI entry points onto traditional operating systems, but instead restructures resource allocation and capability scheduling logic from the ground up, enabling agents to manage hardware and software services much like an operating system kernel. The dual-domain three-step memory structure achieves industry-leading performance in terms of memory persistence and retrieval efficiency (15 milliseconds recall), far surpassing the short-term memory capabilities of traditional assistants. The edge-cloud collaborative routing and cascaded model scheduling (from edge-side 300M-10B to cloud-side 1.5T+) reflect a refined consideration of balancing real-time performance with deep reasoning, a feature that is relatively rare in current personal agent products.

In terms of practical value, StepAmoo targets frequent cross-application, multi-step task scenarios in everyday user activities—such as travel planning, office workflows, and lifestyle services. It aims to free users from the tedious process of repeatedly switching between different applications and performing manual operations, realizing the concept of "intent as a service." If this vision can be consistently realized, it will significantly enhance digital life efficiency. However, its current closed beta phase and specific hardware dependencies limit immediate usability, and ecosystem development (third-party application compatibility with atomic interfaces) also requires time, which means its practical value is still in the verification stage.

StepAmoo's target user base is clearly defined: heavy digital users, business professionals, efficiency seekers, and tech enthusiasts eager to try new products. For these users, the efficiency gains provided by the agent may outweigh the learning costs and ecosystem limitations. General users may need to wait until the product matures and the ecosystem expands before considering it.

In terms of future development potential, StepAmoo's core technical framework (operating system-level agents, dual-domain memory, atomic capability orchestration) is highly scalable. If Step AOS can attract more terminal manufacturers and application developers to join, forming a healthy ecosystem loop, StepAmoo has the potential to become a major benchmark in the personal agent field. However, it must also address challenges such as privacy trust, the initial cold start of the ecosystem, and the cultivation of user habits. Overall, StepAmoo represents a bold and directionally correct technological exploration, and its success will largely depend on the speed of ecosystem development and the establishment of user trust.

7. Application Scenarios

  • One-Stop Smart Travel: Users only need to state their destination and preferences (e.g., "Depart tomorrow morning for Shanghai, with a budget under 2000"), and StepAmoo automatically completes the cross-application closed-loop for flight/high-speed rail inquiries, hotel bookings, ride-hailing, and route navigation. It can also optimize choices based on historical travel habits stored in its dual-domain memory (e.g., preference for window seats, a specific hotel chain) and synchronize the schedule to the calendar.

  • Cross-Application Office Efficiency: Integrates with applications such as WPS, calendar, and email to automatically organize action items from meeting minutes and sync them to the schedule, achieving seamless connectivity from document editing to task distribution. For example, when a user says, "Send today's afternoon meeting minutes to the attendees via email," StepAmoo will automatically extract key points, generate an email draft, attach files, and send the email.

  • Local Life Services: Based on user taste preferences and historical orders, StepAmoo can automatically place food delivery orders on Meituan, plan pickup routes on Gaode, and purchase movie tickets on Maoyan. Users simply need to say, "Order my usual spicy hot pot and have it delivered to the office," and the agent can complete the entire process of searching, placing an order, and making the payment, handling daily life tasks with just one sentence.

  • Content Creation Assistance: Invokes creation tools like Jianying to automatically search for photo materials in the album, generate video drafts, and publish them to Weibo or Douyin. Creators only need to express their ideas (e.g., "Create a 30-second short video from last week's travel photos with upbeat music"), and StepAmoo can complete the entire workflow, including material selection, editing, adding music, and publishing.

  • Personalized Schedule Management: Identifies user habits based on long-term memory and proactively reminds of meetings, recommends today's outfit, and books commuting vehicles at appropriate times. For example, it automatically announces the schedule and weather every morning and calculates the departure time based on the meeting location in the calendar, then calls a ride-hailing service, acting as a 24/7 personal assistant.

8. FAQ

Q: When will StepAmoo be publicly available?
A: StepAmoo is currently in the early co-creation and internal testing phase. The specific dates for public beta or official release have not been announced yet. We recommend following official channels of Jieyue Star (website, Weibo, WeChat official account) for the latest updates. Internal testing access is typically granted through official recruitment activities.

Q: Does StepAmoo require an internet connection to use?
A: StepAmoo supports edge-cloud collaboration. Some simple tasks (such as setting an alarm, local queries) can be completed offline using the on-device model, without requiring an internet connection. However, complex tasks (such as cross-application planning, deep reasoning) require a stable network connection to invoke cloud-based large models (Step Flash/Step Pro). Users can choose the task execution path based on their privacy needs.

Q: How is privacy and security ensured in StepAmoo?
A: StepAmoo features a four-dimensional security governance framework: all operations are based on user authorization (controllable permissions), the execution process is real-time visible (transparent process), sensitive operations require secondary confirmation (reversible behavior), and all operations are logged for audit (trusted operations). Users can manage their authorizations at any time in the settings and revoke access to specific applications or data.

Q: What is the core difference between StepAmoo and other voice assistants (such as Siri, Xiaoi)?
A: The core difference lies in the system identity and task capabilities. StepAmoo is an operating system-level native agent, possessing system-level permissions and capable of directly scheduling hardware and atomic services to achieve cross-application long-chain task closure. Traditional voice assistants are application-level entry points with limited permissions, making it difficult to complete complex, multi-step, cross-application tasks. Additionally, StepAmoo's dual-domain long-term memory capability is something traditional assistants lack.

Q: Which languages does StepAmoo support?
A: Currently, StepAmoo mainly supports Chinese, including Mandarin and common Chinese interaction scenarios. Support for additional languages (such as English) may be expanded in the future based on market demand. Please refer to official announcements for specific details.

Q: Is StepAmoo open source?
A: StepAmoo is a commercial product of Jieyue Star and is not currently open source. The underlying technologies (such as Step AOS, atomic capability engine) are core company assets, and there are no announced plans for open sourcing them yet. However, Jieyue Star has taken open-source actions in the foundational model space (such as certain models in the Step series). Whether the agent system will be open sourced depends on official decisions.

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.