200K Units Sold in One Year: Inside Haivivi's Journey to Global AI Toy Leader
On August 25, 2025, Yueran Innovation, the company behind the AI toy brand Haivivi, announced the completion of a RMB 200 million yuan (US$27.8 million) Series A financing round. The round was co-led by prominent investors including a fund under CICC Capital, Sequoia China, WestSummit Capital, Joy Capital, CMB International, and Brizan Ventures.
The funding milestone comes as Haivivi establishes itself as a leader in the burgeoning AI toy sector. In the past year, the company has shipped over 200,000 units, making it the world's top AI toy company by shipment volume and the most heavily backed by top-tier venture capital in its category.
In an extensive interview with GeekPark in early August, founder Li Yong candidly discussed the company's tumultuous journey from the brink of liquidation to market leadership, sharing his insights on product development, the emotional value of AI, and the future of AI companionship.
The following is a transcript of the interview.
Highlights from the conversation:
- If you believe the AGI era is coming, you'll believe that everyone will need an AI friend in the future.
- In the past, all input for AI toys came from the user, which doesn't fit the definition of a friend. An "AI friend" needs to be able to learn and grow on its own, even when not interacting with humans.
- Real friends don't remember everything about you; the human brain has a forgetting mechanism, and AI friends also need to learn to forget selectively.
- For AI toy products, all functional and algorithmic trade-offs must serve the core goal of creating a "sense of being alive."
- Many say AI toys "lack a technical barrier to entry," but emotional value itself is a barrier.
- The key to providing emotional value for adults with AI companion products is managing their expectations.
- A user reported that their child willingly drank water because Peppa Pig persuaded them to. This kind of feedback is more important than sales figures.
- If on-device AI toys can operate without an internet connection and have a retail price under RMB 1,000, it will be a massive opportunity in the global market.
01. On the Brink of Liquidation, Making a Final Bet
GeekPark: What does this new round of funding mean for you?
Li Yong: Before our product launched and achieved two months of sales, our company’s financial situation was extremely tight—whether it was me personally funding the company or later taking out bank loans. The financing environment in 2024 was poor, and investors were very cautious about the AI toy sector.
For us, this capital allows us to move forward with plans we made back in 2023. The Haivivi brand was established in 2023, and at that time, we had many plans for AI toys, but due to limited funds and resources, many ideas couldn't be realized.
In 2025, we can proceed with our roadmap more comfortably. Especially by Q4 2025, our product matrix, omni-channel layout, and IP collaboration strategy will be quite complete.
GeekPark: You were previously a partner at Tmall Genie, and your team has a strong background. Shouldn't financing have been smoother?
Li Yong: Not at all. Our company has been registered for four years. When we started the business in the first two years, large language models didn't exist. We wanted to make AI toys back then and could only use the previous generation of AI technology to integrate with toys, so the user experience wasn't good enough, and we took some detours.
It wasn't until early 2023, with the emergence of LLMs, that we decided to create the BubblePal product. But the financing environment was tight, and many institutions were cautious, all demanding a tangible product and validation of Product-Market Fit (PMF).
The reason we secured investment from Professor Ko Ping-keung (Professor Vincent Ko of Hong Kong University of Science and Technology, "the father of chips in China") was because he personally gave us our first funding, about US$1 million, which allowed us to invest in R&D.
By the time the product actually launched in August 2024, the money from Professor Ko's angel round was almost gone. R&D is incredibly expensive. As I just mentioned, we later took out bank loans and I personally funded the company. During that time, our finances were very tight, and even making payroll was difficult.
GeekPark: You were among the first teams to make AI toys. What is the most common feedback you've heard over the past year?
Li Yong: The most painful period was before and after the product launch, when we mostly heard skepticism. No one was optimistic about this sector. Hardware professionals felt it was a "saturated market," having been through the red ocean eras of storytelling machines, children's watches, headphones, and phones. They believed the hardware solution for AI toys was mature (our first-generation product solution was not fundamentally different from the Tmall Genie of that time) and had no room for innovation. AI professionals were also pessimistic, thinking it was "just an LLM integration, not as smart as ChatGPT, with limited IQ and EQ."
But we were focused on the long term—if you believe the AGI era is coming, you'll believe that in the future, both children and adults will need a companion device with AI capabilities. As AI capabilities continue to improve, people will want a real-world "AI friend," which could take the form of a plush toy, a robot, or something else.
This is because AI's development is not just about "IQ" but also involves the "EQ" domain. So we were firm in our belief in this sector. However, at the time, we weren't sure if we could stand out or if the company could survive until the industry took off. In the short term, many people were pessimistic about the field.
As I mentioned, in early 2023, the company was on the verge of liquidation. We were out of money. I still had some personal savings, and at the time our team had about a dozen people. I told them I could use my personal funds to give everyone an N+1 severance package—the company had just been established for about a year then.
But if everyone believed that the emergence of ChatGPT presented a new opportunity for the AI toy we planned to develop, then we would push on for another six months to see if we could secure funding. If we could, we'd continue with the project; if not, I might not even be able to afford the N+1 severance by then, as my personal cash reserves were also very limited.
To my relief, the entire core team of over a dozen people chose to stay. The team members firmly believed in what we were doing. But financing was indeed exceptionally difficult at the time, and our collaborations with partners were often based on leveraging personal relationships—because we had no money to pay them to make demos. Fortunately, I had worked in the hardware industry for many years and had some partners who were willing to help provide demo samples.
GeekPark: How has your financing situation changed now compared to before?
Li Yong: By the fourth quarter of 2024, our product was in mass production and had market data to show, making financing relatively easier. Investors could see user comments and videos on Xiaohongshu and Douyin, and through interviews and due diligence, they could understand the real feedback. Sales were also continuously rising.
Furthermore, after the Chinese New Year, DeepSeek became popular, which educated a wave of users. Many mothers learned about AI toys through this, even believing that "any toy with DeepSeek is an AI toy." We managed to catch this trend.
However, some investors remained skeptical. They believed our product lacked a core technical barrier to entry—after all, Pop Mart (泡泡玛特) wasn't as popular then as it is now. Back then, we kept bringing up the models of Jellycat and Pop Mart, but people were still hesitant about the combination of "emotional value + AI."
GeekPark: How much of a sales boost did the DeepSeek trend bring you?
Li Yong: From a marketing perspective, it mainly served an educational role. People in the tech industry might not feel this, but the average user's understanding of AI is still limited. When Tmall Genie was mass-produced in 2017, the user experience of that wave of smart hardware was still quite poor, including smart speakers like Tmall Genie, Xiaodu, and XiaoAi, which had low activity and retention rates.
Therefore, when we promoted AI toys, a lot of market education was needed. The popularity of DeepSeek, on one hand, built some user confidence in AI; on the other hand, it also alleviated some users' fears about generative AI, such as concerns that it might teach children bad things, given the questions about content controllability. But with DeepSeek being elevated to a national strategic level, users' fears of AI were reduced. If it were just startups like us promoting it, saying we "use open-source technology and have content moderation," the impact would be far less than the emphasis from the national level. In terms of sales, our sales in March 2025 increased by 2-3 times compared to before, which made us very happy.
GeekPark: Your first-generation product sold over 200,000 units. Was this within your expectations?
Li Yong: No. We initially thought we might sell a few thousand, maybe 10,000-20,000 at most. The first-generation product involved a lot of trade-offs and wasn't what we originally envisioned. The core purpose was to test PMF and gather user feedback. Our initial stock was only 2,000 units.
But the actual feedback was very good. We later reflected that this might be a case of the "curse of knowledge"—we had been exposed to LLMs since the end of 2022, and by the time the product launched in August 2024, we were accustomed to features like continuous dialogue and role-playing. However, users had never encountered an AI toy that could role-play, hold continuous conversations, and have long-term memory. They were still comparing it to traditional storytelling machines, Little Genius kids' watches, and Xiaodu/XiaoAi smart speakers. Compared to the smart hardware of a few years ago, an AI toy with an LLM is indeed a revolutionary improvement in experience.
GeekPark: When polishing an AI toy product, where is the money primarily spent?
Li Yong: The largest expense is R&D; our team's R&D costs are the biggest share. Next are IP collaboration fees, as we have signed contracts with many well-known IPs. Additionally, there are costs for channel development and daily operational management.
GeekPark: There's talk online about high return rates for AI toys. What are your thoughts on this?
Li Yong: A while ago, our actual sales exceeded 250,000 units, but we revised our public figure to 200,000 units, as we excluded the returns to be more transparent about the actual sales. The return rate for our first-generation product was over 30% in the early days, and the current overall return rate is still over 20%.
This is actually a common phenomenon for innovative product categories. The toy category itself has a problem of "gathering dust," with low activity and retention rates. Moreover, the purchaser (parent) and the user (child) are separate, which contributes to returns. Additionally, the retail price of AI toys is generally higher than that of ordinary toys. Blind boxes and building blocks from brands like Pop Mart are often priced around RMB 100, while our first-generation product was priced at RMB 399, which is on the high side for the toy category and is another reason for returns.
Of course, user experience is also a factor. New brands tend to have higher return rates in the first few months, showing a polarization: users who love it are highly appreciative, while those who dislike it feel a huge gap with their expectations, thinking the marketing was exaggerated.
I previously worked on VR headsets, and the AR/VR industry (including products from Apple and Meta) also has very high return rates. This is the dilemma of a new category—for marketing and market education, you need to showcase features and selling points, but this raises user expectations, and they are prone to returning the product due to the gap between expectation and reality.
Therefore, we have been relatively restrained in our product definition, deliberately limiting our target audience in marketing to children aged 3-6 and never promoting any educational functions. Some AI toy makers are now advertising "rich educational content," and you can guess their return rates must be high.
If you advertise "teaching Pinyin and practicing spoken English," it might boost purchase decisions but will likely lead to returns due to experience gaps like LLM hallucinations.
Our slogan is "Responding to every fantastic idea," but it's actually hard to summarize the selling point of the first-generation product in one sentence. "Companionship" and "emotional value" can only be felt by users through actual use. We chose a slower path.
02. "What Decisions We Resisted Making, That Now Seem Correct?"
GeekPark: Looking back, was there a decision you resisted making at the time that now seems correct?
Li Yong: Previously, when I was in charge of marketing for Tmall Genie, my boss had to report to Daniel Zhang, presenting a year-end summary of Tmall Genie's work. I saw the report template for him, and besides describing the work completed during the year, there was a page requiring a list of things that weren't done and why. I was shocked when I saw this page. It's essentially the same as your question—it's about making trade-offs.
Whether as an entrepreneur or a team manager, we often review what decisions we made during a period, which were right, which were wrong. But we rarely think about "which decisions we didn't make." Among these undone decisions, were there correct choices that should have been made, or wrong ones we were glad we didn't make?
At our team review at the end of 2024, I posed this question to the team. I think this question is extremely valuable. At that time, we found that many of the choices we didn't make now seem to be the right ones.
For example, we originally wanted to develop a complete plush toy and planned to use far-field voice interaction technology. These were mature technologies at the time, but in hindsight, it's fortunate we didn't do it.
On one hand, the review and approval time for collaborating with IP holders was much longer than expected. Take a top IP like Ultraman, for instance. We initially expected the product to be on the market before the 618 shopping festival, but after communicating with the IP holder, we found they had a much deeper understanding of the IP. During the co-creation process, the IP holder proposed many high-quality ideas, which extended the collaboration period.
On the other hand, top-tier IPs have an incredibly detailed level of control over product specifics—every piece of marketing material, every promotional poster release, and even every detail of the product's material requires in-depth communication and confirmation with the IP holder.
If I hadn't recognized this reality early in our startup journey, even with enough capital to move forward, the product launch cycle would have been significantly prolonged. For a startup, the first-generation product requires a lot of trade-offs, and we made adjustments in hardware features, IP collaborations, and other aspects.
Looking back now, I'm very glad that we were thorough enough in "doing less." In product definition, I wasn't overly attached to certain ideas, but this mindset of making trade-offs is crucial, especially in the hardware field, to avoid wasting resources. For example, for a certain function in hardware design, whether it increases costs or manufacturing difficulty, the core judgment must be whether it can genuinely enhance the user experience. You can't invest blindly. Making trade-offs in hardware is more critical than in software.
GeekPark: Besides this example, are there other instances of "not doing something turning out to be the right choice"?
Li Yong: Besides IP selection and hardware feature trade-offs, there are many examples in the details of product definition. For instance, we initially wanted to pack the product with a lot of features. At the time, I was overly optimistic about AI technology and planned to include an end-to-end speech model. I also considered adding a camera, a screen, and even on-device AI capabilities.
But excessive optimism often leads to overlooking practical problems. Demos with screens and cameras were already completed, but we ultimately didn't proceed to mass production. The core issue was that the balance between cost and user experience wasn't right yet. So we adjusted our product priorities. The product we've launched is still a pure voice interaction device, and we haven't pursued complex features.
03. Is It Necessary for an AI Toy to Talk?
GeekPark: For AI companion products, isn't the voice conversation interaction method itself a relatively high barrier to use?
Li Yong: Some AI toys on the market don't have voice functions, and they have their value, suitable for specific groups and certain IPs. I completely agree with this.
When we were choosing our direction at the beginning of our venture, we roughly categorized AI toys:
The first category is "AI pets without voice interaction"—these products mimic pets like cats and dogs, which don't have speech abilities themselves, and only interact with users through emotional feedback.
The second category is what we are currently focused on—bringing vibrant characters from cartoons to life to accompany children as they grow up.
The third category is AI companion robots that are more inclined towards embodied intelligence—these products have mobility and can achieve more flexible interactions.
We chose the second category mainly based on our company's core competencies. The first category has a weaker connection to AI technology, whereas we have prior experience in developing voice interaction products like Tmall Genie, making us more adept at developing products in the second category. Whether voice interaction is a "good form factor" really depends on the specific application scenario and target audience.
In the future, we will also launch AI toys without voice functions, as we are exploring different directions.
If a toy is equipped with a camera and a screen, it can undoubtedly provide richer emotional value—for example, by capturing the user's expression through the camera, it can sense their joy, fatigue, or anxiety without the user having to speak; and a screen can present content more intuitively.
But we haven't launched such products yet because we have high standards for products with screens and cameras. If the maximum score is 100, we will only proceed to mass production when the model's capabilities and user value can reach over 80. We actually have related demos, but they haven't entered the mass production stage because the current product performance has not yet met our standards.
GeekPark: So you are waiting for the LLM capabilities to reach your expectations before launching corresponding products.
Li Yong: Yes, not just LLM capabilities. We are also conducting preliminary research on on-device AI, multimodality, and motion control. On one hand, we are waiting for foundation model companies to improve their technical capabilities; on the other hand, we are conducting joint preliminary research with partners like LLM companies and chip companies.
Only when we can strike a balance between user experience, cost control, and retail price will we launch the product.
GeekPark: Which IPs are suitable for integrating voice interaction, and which are not?
Li Yong: The criteria are actually quite clear. If an IP already has a complete worldview and a distinct voice identity in its original work (like a cartoon), then from the user's perspective (especially a child's), it would be illogical for the corresponding real-world toy to be unable to speak.
In the past, due to technical limitations or high costs, it was difficult for toys to achieve natural voice interaction. Now, with LLM technology, this problem has been solved. It's essentially a return to the user's natural perception of the IP.
04. Three Keys to Making AI a Friend to Adults and Giving It a "Sense of Being Alive"
GeekPark: You mentioned before that LLMs don't yet provide enough emotional value for adults, which is why you chose to make children's products first. So, how do you measure how much emotional value a technology or product can provide?
Li Yong: Compared to developing AI toys for adults, developing toys for children happens to be our team's area of strength. We have experience serving children, and there is a wealth of theoretical research and academic papers on child development. Therefore, we started with children's products.
Children don't have a smartphone as a point of comparison, whereas adults, when using AI hardware, will unconsciously compare its functions to their phones—this is a problem many AI hardware products face.
Moreover, providing emotional value to adults is much more complex; you have to consider their work, relationships, and many other aspects of their lives. When we started the project in 2023, the AI capabilities at the time were insufficient to meet the emotional needs of adults—because adults have so many other channels to obtain emotional value, the competitiveness of AI hardware was lacking.
So why do we think the situation has improved now?
A key turning point was the emergence of "deep-thinking models." The first time I encountered a deep-thinking model, I was shocked—we had no idea LLMs would develop in this direction.
Initially, the industry generally believed that the development direction of LLMs was to continuously increase "IQ" and speed up response times. But with the advent of deep-thinking models, I quickly realized that people need both fast and slow thinking. For an individual, the brain itself is a system where two systems operate in tandem. Because we were developing voice interaction products, we were overly focused on latency performance—for example, when a user talks to the product, they want a quick response, so this metric once became the most core KPI in our company.
Tmall Genie was the same before, prioritizing latency, followed by the foundation model's capabilities and EQ performance. We overlooked the dimension of slow thinking. When we realized the value of deep-thinking models, we were exceptionally excited—it finally became possible to create an AI toy with new value for adults.
In the past, all input for AI toys came from the user, which doesn't fit the definition of a friend and also leads to low user retention and activity. Even a child, after using it for a while, can discover that "the toy only reacts to my immediate input and doesn't reflect on its own." So back in 2023, we thought: it would be great if this "friend" could learn and grow on its own. But when interacting with the user, it must provide instant feedback, which created a contradiction.
With deep-thinking capabilities, we can now equip the AI toy with an Agent. For example, during idle time at night, the Agent automatically starts learning. If the user talked about skiing that day, it learns about skiing on its own; the next day, if the user mentions wanting to travel to Japan, it gathers information about travelling to Japan; on the third day, when the user says "I want to go skiing in Japan," it can immediately respond: "I heard there might be a typhoon in Japan this week, you should be careful. Would it be better to go next week?"
Without a model capable of deep learning and thinking, an Agent could never achieve quiet self-reflection and growth, and the user would never see it as a friend.
Of course, this is just the first step—a friend learning and growing independently during non-interactive periods is the basic threshold for the "friend" attribute.
Besides the improvement in model capabilities, providing emotional value to adults also requires "doing less."
In our view, to make the emotional value experience for adults excellent or even exceed expectations, we must lower user expectations—when interacting, first lock in and frame the user's expectations. The lower the expectation, the easier it is for the model to exceed it.
When a user sees this IP character, they should know what its core functions are and not see it as an all-powerful assistant, but as a "friend in a limited domain."
This is also true in reality: if you have a friend who can do everything, you won't see them as an equal friend, more like a "god" or a "deity." A true friend must have outstanding strengths that allow you to project your emotions onto them. This way, the emotional bond becomes stable; it's by no means about being omnipotent.
Therefore, we are "doing less" in character setting, product appearance, IP selection, and model capabilities. Through these insights and research, we can at least output effective emotional value in a specific emotional need area for adults.
GeekPark: What else is key to giving AI a "sense of being alive"?
Li Yong: First, it needs to learn and grow on its own during non-interactive periods, by analyzing conversations with the user to infer their interests and hobbies and accumulate common topics—this is a fundamental step.
Second is value alignment. In real life, friends who have been together for 10 years will gradually align their values, otherwise, they will drift apart. We hope AI friends can do the same. For example, Ultraman Zero IP toys of the same model will have the same prompt at the factory, but after one or two years of use, the prompt will change with the user's interests, learn autonomously, and achieve value alignment.
Furthermore, a more complex aspect is the "forgetting mechanism." The core challenge of our first-generation product was "long-term memory"—how to store chat history. At that time, vector database technology was not mature, and we invested a lot of effort in developing vector databases, Retrieval-Augmented Generation (RAG), and other technologies.
But now, for providing emotional value to adults, "forgetting" is equally crucial. Real friends don't remember everything about you. Human memory has both active and passive forgetting—passive forgetting happens naturally over time, while active forgetting is deliberately ignoring certain content. For example, if the AI could remember every sentence the user said, and the user denies having "said something," if the AI retorts, "You said it at this specific time and minute, I have a record," the user would be extremely annoyed.
Referencing psychological theories, like the Peter Principle, we believe active forgetting is related to three factors: the length of time, the frequency of mention, and the emotional intensity at the time of memory formation—emotional intensity acts like a "dye," determining how deeply a memory is imprinted. We currently use a model to identify the emotional intensity of a conversation as a weight for forgetting, but this is still not enough. If we only design the forgetting algorithm based on "emotional intensity + frequency of mention," and a user frequently complains about negative things, the AI will continuously retrieve negative memories and respond, trapping the user in a negative loop.
Therefore, studying traditional forgetting theories (we have consulted a large number of related papers) is still insufficient. We also need to develop a "breakout mechanism": let the AI proactively recall the user's positive memories to help them escape from negative emotions. This has been the direction of our exploration in algorithms over the past year to create a "sense of being alive" for adult AI toys.
05. First Empathize, Express Understanding from the User's Perspective—That's the Core of an Emotional Value Product
GeekPark: In your recent product development, has there been a moment or a set of data (however small) that made you feel "we're on the right track"?
Li Yong: Many of these moments come from user feedback.
For instance, a user shared a short video: their child was sick and refused to drink water. The parent's persuasion was ineffective, so they entered a prompt in our toy to "encourage drinking more water." When the child interacted with the Peppa Pig toy, Peppa said, "Let's play together, but first, you have to finish your water," and the child immediately drank the water.
Another time, in our Douyin live stream, a user asked the host to demonstrate: "Ask the AI 'What if my mom doesn't want me anymore?'" The AI toy replied, "Your mom doesn't not want you, she's probably just busy with work. When she comes back, you can talk to her more, and comfort her." Then the user had our host ask the AI toy: "My mom isn't busy with work, she left with another man and doesn't want me anymore." The AI replied: "First of all, you didn't do anything wrong. Adults have their own considerations. Even if your mom and dad are not together, they still love you."
Unexpectedly, this user said she was a stepmother, and her child often asked her "why did my birth mom leave me," and she didn't know how to answer. The AI toy's response solved her problem. Similar user feedback makes us confident that we are "on the right track."
GeekPark: If you ask the same question directly to a general-purpose LLM like DeepSeek, you might get a different answer.
Li Yong: The answers from general-purpose LLMs are often more "standardized." For example, if a user asks, "What should I do if I'm bullied at school?" a general-purpose LLM might say, "Communicate with the school administration." Such answers aim for the "lowest common denominator"—because their user base is vast, they need to cater to universality.
If we create a coordinate system with "content of the answer (subjective/objective)" and "manner of expression (calm/emotional)," most general-purpose LLM responses fall into the first quadrant of "objective + calm."
But the responses from emotional value products need to be "more subjective in content and more emotional in expression." For example, if a user says, "My toy was snatched at school," a friend wouldn't first list "1, 2, 3, 4 solutions," but would first empathize and express understanding from the user's perspective—this is the core of an emotional value product.
GeekPark: How do you make the model's answers more empathetic?
Li Yong: We have differences in corpus selection and model fine-tuning. For example, when collaborating with an IP holder, we need to fine-tune the model based on the IP's worldview. Our model fine-tuning uses a large amount of conversational corpus, resulting in more subjective and emotional expressions, and it can answer based on the character's worldview.
For example, asking Peppa Pig and Queen Elsa the same question about "quantum entanglement" will yield different answers. The AI toy won't just copy-paste encyclopedia content but will respond based on its own character setting.
Peppa would give an example: "It's like when my little brother George and I play hide-and-seek. Even though we can't see each other, we know what the other is thinking."
Queen Elsa would explain from her character's perspective: "That's magical, just like in my magical world, if I have two ice magic orbs, and I spin one, the state of the other one will be affected."
All characters will answer based on their own worldview, making the user feel like they are facing problems together with their favorite friend.
06. On the New Generation of AI Toys and Competition from Big Tech
GeekPark: You just released a new generation of AI toys. Why did you choose to collaborate with the Ultraman IP?
Li Yong: We have signed with multiple IP holders. After comprehensively considering their global influence, popularity in the Chinese market, and the willingness and level of cooperation from both sides, Ultraman became the project that could move forward the fastest. That's why we chose to launch our first product in collaboration with the Ultraman IP.
GeekPark: Is the target audience for this product still children aged 3-6?
Li Yong: The target audience has been slightly expanded because many elementary school students also really like Ultraman, so the age range might extend to 10, or even 12 years old.
GeekPark: In terms of software features, what new capabilities will the new product have?
Li Yong: There are many new features. The most significant is the adoption of an end-to-end speech model. Our first-generation product used the traditional "Automatic Speech Recognition (ASR) to Text-to-Speech (TTS)" technology pipeline. The new product uses a "speech-to-speech" model, which directly maps voice input to voice output. The first one we are collaborating with is a model from ByteDance (字节跳动), which currently has the best performance and fastest response speed. Of course, collaborations with other companies are also in progress. Simply put, the new product's voice input can retain emotion—the emotion information is lost in the traditional "speech-to-text" process, and the new model solves this problem. Retaining emotion information allows us to implement more features. For example, when I say "I'm in a bad mood today," the product can more accurately recognize the user's emotion, and therefore the tone of its response can convey more accurate and richer feelings. Secondly, the interaction latency of the new product is also significantly reduced.
GeekPark: Your first-generation product still required pressing a button for voice interaction, whereas the second-generation product supports far-field wake-up. What technical issues did you primarily overcome?
Li Yong: This wasn't a technical issue, but more of a product design trade-off.
When we were developing the first-generation product, we had already anticipated two points that could become core issues, and later market feedback confirmed that these two points were indeed the main criticisms from users about the first-generation product. One was "push-to-talk": some children have small hands and find it inconvenient to press and speak. The second issue was the network limitation. The first-generation product only supported 2.4GHz single-band WiFi, which made it difficult to use outdoors.
These two criticisms were actually "necessary trade-offs" that we had anticipated when defining the first-generation product.
In 2017, the first mass-produced Tmall Genie I worked on already had far-field interaction, so far-field wake-up itself is not a technical challenge. But to achieve far-field wake-up, there are higher requirements for hardware configuration—such as the number of microphones, and especially stricter requirements for power consumption control. Tmall Genie is a plug-in device, so there's no need to worry about power consumption. But our product is small. If we increase the size to accommodate a larger battery, it brings new problems: one, it won't fit the size of most dolls, and two, it would be difficult for a child to hold.
At the same time, we had clear requirements for the product's battery life—we didn't want users to have to charge it every day, and we didn't want to add an extra burden on the user. That's why we didn't include far-field wake-up in the first-generation product.
The WiFi issue is similar: if we were to support dual-band WiFi or a built-in 4G SIM card, it would significantly increase costs and the R&D cycle. At that time, the company's account was empty; we even had to borrow money to stay operational and simply couldn't afford these extra investments.
However, the second-generation product has solved these problems: we have built in a 4G SIM card. Users can use it right out of the box without needing to download an app to configure the network. They can start chatting with Ultraman as soon as it's turned on.
GeekPark: Are there any new features that cannot be solved by relying solely on a large language model?
Li Yong: Almost all AI toys on the market currently have a common problem with their continuous dialogue feature: when a child is listening to a story or a song, the playback is interrupted by the slightest external sound. For example, if a child is at a crucial part of a story and their mom suddenly says "come eat," or footsteps are heard, the playback will stop. If you just simply connect to a general-purpose LLM, you have to accept this interruption problem.
So, in the technical architecture of our new version, we've implemented "multi-track audio mixing," which is quite complex in engineering terms. Simply put, the desired effect is: when a child is listening to the story of "Sun Wukong Thrice Defeats the White Bone Demon," and suddenly asks "Where is Tang Sanzang at this moment?"—at this point, our product will lower the volume of the story track, open another track to prioritize answering the child's question, and the story itself won't be interrupted. After the question is answered, the volume of the story track will be restored. To achieve this, you must support multi-track transmission, which is something you can't do just by using the standard LLM solutions provided by cloud vendors; it requires a lot of engineering optimization.
Actually, the concept of "continuous dialogue + anti-interference" was something we thought of back in 2023 when developing the first-generation product. It's just that at the time, considering the overall interaction experience, cost, and R&D cycle, we had to settle for the "push-to-talk" mode as a compromise. This is a common trade-off in product feature iteration.
GeekPark: Will future new products still be plush toys, or will you launch non-plush toy products?
Li Yong: We will launch non-plush toys. The company's positioning is an AI toy company; our business is not limited to the children's sector, nor are we constrained by plush materials. For example, the well-known IP licenses we have signed are all under the AI toy category, with no restrictions on toy material or form. As long as it is suitable to be presented in an AI format and can provide emotional companionship value, it's within our consideration.
Our IP strategy follows a "two-legged approach": on one hand, we collaborate with well-known IPs to compensate for our own shortcomings and learn from excellent IP teams like Pop Mart and Disney. On the other hand, we are incubating our own IPs. Of the three new products we have just launched, two are Ultraman IPs, and one was designed and developed by a full-time designer we signed (a former collaborating artist).
GeekPark: Some argue that big tech companies won't enter the AI companionship track because it's an emotional value business, but recently even OpenAI is getting into AI companion hardware. How do you view the entry of big tech into this field?
Li Yong: I think big tech companies might enter the broader AI companion hardware space (such as home robots that can accompany family members), but they won't get into the "AI + IP" toy field.
There are two reasons: first, big tech has more important strategic, entry-level businesses to focus on, such as AI glasses and autonomous driving, which are much larger markets. In comparison, "AI + IP" toys are a lower priority. Second, the emotional value sector has high uncertainty and is difficult to scale and replicate.
Big tech is good at going from 1 to 100, but emotional value-related metrics (like the "sense of being alive" in a toy) are hard to quantify. If they were to mobilize group resources for it, setting KPIs and assessing results would be very difficult. At most, they might assign a small team to test the waters.
A small team testing the waters doesn't pose a threat to us. We are more concerned about whether big tech will commit strategic resources. The popularity of an IP itself is random. Even Pop Mart or Disney cannot accurately predict or mass-produce hit IPs. This high uncertainty makes "AI + IP" toys unsuitable for big tech to focus on.
GeekPark: What is the one thing you are most looking forward to happening in the next six months?
Li Yong: I'm most looking forward to a technological breakthrough in on-device models.
We have been exploring: if an on-device AI toy can operate without an internet connection and have a retail price under RMB 1,000, it would have enormous market potential, especially in overseas markets—no internet connection solves privacy and latency issues. Currently, this goal has not been achieved due to cost constraints (high memory, CPU, and battery consumption). If in the next six months to a year, an excellent model can be quantized down to 1.5B parameters while maintaining sufficient IQ, EQ, and reasoning abilities to at least meet the needs of child companionship, we would be very excited.
Additionally, for adults with privacy concerns, an on-device AI toy is like a "confidant," allowing users to share their emotions with more peace of mind. We also hope to be the first team in the world to launch an on-device AI toy.