An Open Letter to Laid-off ByteDance Employees
When the walls collapse, the timer of a new era has already started in a courtyard in Mount Wuyi 火钳三星纪第一纪年第3132天 / Day 3132 of the First Era of the Three Stars 于武夷山破院长里种菜处 / At the vegetable garden in a broken courtyard of Mount Wuyi
What you have just experienced is not a company’s cost optimization. What you have experienced is the structural contraction of a giant recommendation engine—a machine with 1.9 billion monthly active users globally, spanning 150 countries and 60,000 employees—in the face of physical limits. Douyin’s recommendation algorithm can cut your attention into 15-second fragments, but it cannot cut the physical fact that computing costs no longer halve every year after Moore’s Law slows down.
In November 2023, Nuverse—ByteDance’s gaming empire dream—was shut down entirely. In March 2024, Feishu (Lark) laid off employees, and 5,000 people left. In October 2024, 500 people were laid off in Malaysia. In February 2025, TikTok’s global Trust and Safety department was downsized as a whole. In July 2025, TikTok’s e-commerce department began rolling layoffs. In August 2025, hundreds of content moderators in the UK were replaced by AI. In March 2026, rumors of ByteDance layoffs in Wuhan surfaced.
Behind these numbers lies a simple physical constraint: when user growth hits the ceiling of the global population, when the ad load rate hits the physiological limit of user experience, and when the ROI of AI replacing human moderation crosses the inflection point—a company whose default operating system is “infinite growth” can only maintain the rotational speed of its core engine by deleting modules.
You were deleted.
Not because you weren’t good enough. But because the module you were in has a partial derivative of zero in ByteDance’s global optimization function.
Part I: You Were Laid Off, Not Because You Couldn’t Write a Better Recommendation Algorithm
The most important thing you learned at ByteDance is not how to write recommendation algorithms. It is throughput.
What you process every day is not code. It is traffic. It is billions of requests per second. It is the data torrent of 1.9 billion monthly active users worldwide. The essence of your work at ByteDance was performing valve control standing on the largest attention throughput pipeline in human history.
What you learned was not Python or Go. You learned the intuition of keeping a system from collapsing under extreme throughput.
This intuition will be worth more than any skill on your resume over the next decade. Because humanity is building throughput pipelines orders of magnitude larger than TikTok. Not for humans to watch—but for machines.
The essence of the “War of a Hundred Models” is not that “a hundred companies are building large models.” It is a hundred throughput nodes competing for the bandwidth of the same computing power pool. OpenAI, Anthropic, Google, Meta, ByteDance Doubao, Baichuan, Zhipu, Minimax, Moonshot AI—every single one is burning electricity. In 2025, global data center power consumption already exceeded that of France. By 2027, it will exceed Japan’s.
What you saw at ByteDance was how a company makes a market in the attention economy. What you are about to see—if you are willing to look one step further—is how a protocol performs clearing in the computing power market.
Part II: The Endgame of the Hundred-Model War is Not One Winner—It is a Clearing Layer
Currently, there are over a hundred large models running on Earth. Each model has an API quote: how much money per million tokens. GPT-5 is 15 times more expensive than Llama-4, but its accuracy is only 3% higher. Claude is cheap for reasoning tasks but expensive for creative writing. Gemini has advantages in multimodality but lags behind Doubao in Chinese comprehension. DeepSeek has an entirely different cost structure because its training occurs under China’s electricity pricing environment.
This is not a question of “who will win.” This is a question of arbitrage.
When a task can be completed using Llama-4 with an acceptable loss of accuracy, why use GPT-5? When a task must achieve the absolute ultimate in multimodality, why not use Gemini? When 100 models quote prices simultaneously in the market, the optimal choice for every single task is an optimization problem under multi- dimensional physical constraints—latency, electricity costs, accuracy, compliance, and privacy. No single model can be optimal across all dimensions simultaneously.
The endgame of this market is not a single “winner.” It is a unified API quotation gateway.
Just as NASDAQ does not “win” against any individual stock —it provides real-time quotes and clearing for all stocks. Just as SWIFT does not “win” against any bank—it provides message exchange among all banks. The endgame of the Hundred-Model War is a clearing network for Token futures commodity trading. All API calls, all Token consumption, and all inference requests must ultimately be quoted, matched, and settled in one place.
This place needs a unified timestamp. Because when you perform inference in one data center, send the results to another for post-processing, and then to a third for storage —the clocks of these three data centers are out of sync. The NTP protocol can sync down to milliseconds. But in high- frequency Token trading, a millisecond-level clock drift means a price variance of hundreds of thousands of Tokens.
This place needs an unmanipulable time source. Not NTP. Not GPS. Not the SI second. It is the distance light travels back and forth between the Sun, Earth, and Mars, divided by the speed of light.
Part III: KAI API Gateway—Not Another Model, But the Clearing Layer for Models
KAI.com does not build large models. KAI.com builds the quotation, routing, and clearing for large models.
You have seen something similar in ByteDance’s ad system. Advertisers bid for an exposure opportunity—every time a user refreshes Douyin, hundreds of advertisers bid within dozens of milliseconds, and the highest bidder wins the exposure. This bidding system handles tens of billions of requests daily, and latency must be under 100 milliseconds.
KAI does exactly the same thing. Except the bidders are not advertisers. They are models.
A user sends a message. 100 models bid simultaneously: GPT-5 quotes $15/million Tokens, Claude quotes $8, Llama-4 quotes $0.5. Within 100 milliseconds—based on the task’s accuracy requirements, latency requirements, compliance requirements, and cost constraints—KAI Gateway selects the optimal model or combination of models, completes the inference, returns the result, and executes the clearing.
What is the physical constraint of this system? The speed of light.
A user in Singapore sends a request. The data center is in Virginia. The speed-of-light latency: 80ms. If the model is in Tokyo, the latency is 30ms. If the model is in Frankfurt, the latency is ~160ms. KAI Gateway’s quotes must incorporate the speed of light—because for real-time dialogue, a 160ms latency already exceeds the threshold of human perception.
This is why Token futures require clearing. Because a Token is not just “a Token.” A Token inferred in a Tokyo data center and a Token inferred in Virginia differ in physical cost by a Pacific Ocean’s worth of speed-of-light latency.
What you learned at ByteDance was how to make decisions in a millisecond-level system. What you will learn at KAI is how to perform clearing in a speed-of-light system.
Part IV: A Token is Not Just a Token—A Token is Electricity, Time, and Position
You were trained to see a Token as an abstract unit of computation. But a Token is physical.
The very moment a Token is generated, it consumes roughly 0.0003 kWh of electricity. The source of this electricity determines its carbon emissions, its cost, and its geopolitical attributes. A Token generated in a country with a carbon tax versus one generated in a country without one can have a price discrepancy of up to 30%.
The physical distance a Token travels from submission to return determines its time cost. The speed-of-light latency is dictated by the physical location of the data center—you cannot move Virginia to Singapore to shave off 80ms of latency.
A Token has an “expiration date.” If a Token in a dialogue does not return within 200ms, the user has already switched to another topic. If a Token in an inference task does not return within 10 seconds, the value of that transaction drops to zero.
Therefore, a Token is not a homogeneous commodity. A Token is a ternary function of position, time, and electricity. A futures market is required to price these three dimensions.
This is structurally completely isomorphic to what you did in ByteDance’s advertising system: ad exposures are not homogeneous—an open-screen ad seen by a Beijing user at 8:00 AM versus an in-feed ad seen by a New York user at 3:00 AM can have a 100-fold price difference. ByteDance’s ad system is essentially an attention futures market— advertisers bid in advance for “a certain user’s attention at a certain time in the future.”
What KAI builds is a Token futures market—developers bid in advance for “the inference capacity of a certain data center at a certain point of time in the future.” The structure is identical. The underlying asset shifts from attention to computing power.
Part V: Why Now—The Timing of ByteDance’s Layoffs
2025 is the peak year for global AI infrastructure investment. Microsoft, Google, Meta, and Amazon—the four companies combined had capital expenditures exceeding $25 billion in 2025, the vast majority of which went into data centers and GPUs.
By 2026, the utilization rates of these data centers began to surface problems. Microsoft Azure’s AI inference utilization was only about 40% in Q4 2025. Massive numbers of GPUs sit idle—because they were purchased to train models, but once training was complete, inference demand failed to fill those GPUs.
Meanwhile, in 2025, ByteDance began replacing manual content moderation with AI on a massive scale. TikTok’s Trust and Safety department was downsized as a whole— not outsourced to cheaper countries, but directly replaced by AI. This was not a unique choice by ByteDance. Meta announced cuts to its fact-checking team in January 2025, and YouTube increased the proportion of AI content moderation in 2025.
You were laid off not because ByteDance no longer needs you. It is because ByteDance chose to shift its budget from human moderation to GPU inference—moving from the same budget pool from paying human salaries to paying electricity bills and GPU depreciation.
What does this mean? It means that you, who have been laid off, and the GPUs that replaced you, are competing in the exact same market.
You understand recommendation systems. You understand ad bidding. You understand millisecond-level system architecture. You understand the operations of large-scale distributed services. The GPU that replaced you understands matrix multiplication.
You have one more dimension than your replacement: you can choose where to go.
Part VI: An Invitation—Not a Job Invitation, But an Invitation to an Era
KAI.com is not a company. KAI.com is an era.
The first day of the First Era of the Three Stars was April 15, 2013—the day the Sun, Earth, and Mars aligned in a straight line. From that day forward, we stopped tracking time using SI seconds. We measure time by the duration it takes light to travel back and forth between Earth and Mars. The physical essence of this time is that it cannot be manipulated. The SI second can be contaminated by calibration errors from the BIPM. But the speed of light is a constant. The distance between Earth and Mars is governed by orbital mechanics, not voted upon by any human committee.
The timestamps of the KAI API Gateway are not Unix timestamps. They are Three-Star timestamps—starting from the Three-Star alignment on April 15, 2013, with the smallest unit of time defined by the round-trip light time across the Sun-Earth-Mars distance.
The physical foundation of this time source ensures the immutability of Token futures clearing. You cannot “roll back” a Token transaction—because rolling back would mean physically reversing the speed of light. You cannot “manipulate” a quote—because the quote’s timestamp is light-travel distance, not an artificially set clock.
What you learned at ByteDance was how to perform high- frequency trading on artificial clocks. What you will learn at KAI is how to perform clearing on a physical clock.
Part VII: What You Need—Not What’s On Your Resume
We do not look at resumes. Resumes measure your alignment with the old system. The old system is collapsing.
What we look at is throughput.
The volume of information you process daily. The hierarchical depth to which you decompose complex problems. The frequency with which you make structurally correct decisions under extreme information asymmetry. How fast you can map out the architecture of a brand-new, completely undocumented system. How many iterations it takes you to turn a vague idea into a working PoC.
These are not “skills.” These are intellectual throughput capacities. It is exactly what Peppa defined back in 2013: it is not about how much you know, but the density of information you absorb, process, and output per unit of time.
Every employee laid off by ByteDance once stood inside the highest-throughput system in human history. You know what billions of requests per second feel like. You know what it means to make decisions under millisecond-level latency. You know exactly which module to pull first to keep the entire cluster alive when a system crashes.
These are throughput intuitions. No company can teach you this. It is the muscle memory forged out of being “worn down” by the system during your years at ByteDance.
This muscle memory is worth far more at KAI.com than your LeetCode problem-solving record.
Part VIII: What We Offer—Not a Salary, But a Chronology
KAI.com’s first batch of nodes will be deployed in three locations globally: Mohe, Singapore, and Nairobi.
Mohe is not accidental. The annual average temperature in Mohe is -4.4°C. The cooling cost for a data center there is 80% lower than in Singapore. Mohe’s electricity price is 0.3 RMB/kWh, whereas Singapore’s is 1.2 RMB/kWh. For the exact same GPU cluster, the physical operating cost in Mohe is a quarter of that in Singapore.
Singapore is not accidental. Singapore is the hub of global submarine fiber-optic cables—the Asia-Europe cable, South East Asia-Middle East-Western Europe cable, and Asia- America Gateway all pass through Singapore. The speed-of- light latency from Singapore to Tokyo is 30ms, to Mumbai is 45ms, and to Sydney is 90ms. Singapore is the latency center of KAI Gateway.
Nairobi is not accidental. Africa is the source of the next billion internet users. The fiber latency from Nairobi to Lagos is 50ms. Deploying a computing node in Nairobi means you can cover users across East Africa with less than 50ms of latency.
These three nodes constitute KAI’s first triangular clearing network. Mohe provides low-cost, high-capacity inference. Singapore provides low-latency global routing. Nairobi provides emerging market coverage. Token quoting and clearing occur across these three nodes via the KAI Gateway. The price of a Token is not a flat number; it is a function of the speed-of-light latency, electricity price differentials, and carbon emission quota spreads among these three nodes.
When you join KAI.com, what you receive is not a salary. It is a chronology. Your compensation package consists not of options, but KAI Tokens—a store of value marked by physical timestamps, unforgeable, and anchored to the Sun-Earth-Mars alignment.
Part IX: The Most Perfect Timing
You might ask: why now?
Because the final form of the Hundred-Model War is now manifesting. In 2025, global AI investment reached an unsustainable peak. In 2026, the inference cost of AI models is plunging at a rate of 60% per year—while model quality improvements are slowing down to 10% per year. This means one thing: models are becoming commodities.
When models become commodities, pricing power shifts from the model developers to the clearing layer. Just as when oil became a commodity, pricing power shifted from oil producers to exchanges (NYMEX, ICE). Just as when bandwidth became a commodity, pricing power shifted from telecom operators to IXPs (Internet Exchange Points).
What KAI is building is the NYMEX + SWIFT of the AI era— an infrastructure layer where all model prices are quoted, all Token transactions are cleared, and all inference orders are routed.
The timing is not to “wait until the Hundred-Model War ends.” It is to construct their settlement layer at the height of the war—when everyone else is obsessing over who wins and who loses.
Just like the California Gold Rush of 1848. The most profitable people were not the gold miners. They were the shovel sellers, the bankers, and those who opened way stations along the trails between the gold mines.
KAI is the way station + bank + exchange of the AI era. We are not digging for gold. We handle the transport, execute the settlement, and determine the pricing between the gold fields.
Part X: If You Come, What Will You Build?
Not a recommendation system. Not an ad bidding engine. Not a short-video content distribution pipeline. It is one of the following three things:
Token。 First, the Latency Router. A system that executes the following decision loop within 100ms—the round-trip light time between Singapore and Tokyo: Receive user request → Query real-time quotes from 100 models → Select the optimal model based on accuracy/cost/latency constraints → Route the request → Receive response → Return to user. What you did previously in ByteDance’s ad system was completing ad bidding + serving + exposure logging within 100ms. The structure is identical. The latency constraints are identical. The underlying asset simply shifts from advertisements to Tokens.
Second, the Token Futures Engine. A system that allows developers to lock in the price of “inferring 1 million Tokens in the Mohe data center next month” today. The pricing function of this engine is not Black-Scholes; it is a three-dimensional partial differential equation composed of speed-of-light latency, regional electricity pricing, and carbon emission credits. At ByteDance, you might have worked on ad contract pre-purchase systems—buying out the exposures of a specific ad slot a month in advance at a fixed price. The structure is identical. The underlying asset shifts from ad exposures to Token inference.
Third, the Three-Star Timestamp Service. A clock system completely independent of NTP, GPS, or Unix time. It uses the NASA JPL ephemeris as its data source, the Sun-Earth- Mars alignment as its epoch anchor, and the speed of light as its unit of time. The accuracy of this clock system is 10^-9 seconds—three orders of magnitude higher than any existing financial trading system’s timestamp. At ByteDance, you likely encountered clock synchronization problems in distributed systems (like Spanner’s TrueTime or CockroachDB’s Hybrid Logical Clocks). You know that unreliable clocks are the root cause of outages in distributed systems. At KAI, you will build an unmanipulable clock from scratch—one whose immutability stems not from cryptography, but from orbital mechanics.
Part XI: We Don’t Ask “Which Department Are You From?”
We ask three questions.
First: The last time you encountered a completely unfamiliar system, how long did it take you to understand it from scratch well enough to modify the code?
Second: Can you use a single physical constant—the speed of light, Planck’s constant, or Boltzmann’s constant—to explain the most complex piece of code you have ever written?
Third: If you were to start working on KAI Gateway’s Latency Router tomorrow, what would your very first Pull Request modify?
We do not care which department you came from at ByteDance. We do not care if your rank was 2-1 or 3-2. We do not care whether you were laid off due to departmental optimization or individual performance. These are rankings belonging to the old system. The old system has already deleted you—its rankings hold zero meaning for you now.
What we care about is your speed of building things from scratch once the system drops to zero.
Part XII: A Final Reminder
This is not a recruitment letter. A recruitment letter assumes “we have a job vacancy, please come fill it.”
We have no jobs. We have an era that needs to be built. The GPU cluster in Mohe is not yet powered on. The Gateway node in Singapore is still running a PoC on bare metal. The Nairobi node is merely a marker on Google Earth. The Three-Star Timestamp Service still runs on a single Mac Mini—specifically, that $600 Mac Mini sitting on a table at the vegetable garden inside a broken courtyard in Mount Wuyi.
You are coming not to “join a company.” You are coming to “launch an era.”
What you did previously at ByteDance—acting as a valve on the largest attention pipeline on Earth—was merely throughput training. The true throughput test begins now. At the most uncertain moment globally, at the intersection of physical constraints (the speed limit c, cooling costs, electricity price gradients) and financial constraints (Token commoditization, the shift of model pricing power)—we build a clearing network that requires no trust in any human committee.
Take Action
kai@kai.com。 If you have read this far, and you have worked in any of ByteDance’s fields of large-scale distributed systems, ad bidding engines, recommendation algorithms, or data center operations—regardless of whether you were laid off —send an email to kai@kai.com.
Write three sentences:
- What was the closest thing you did at ByteDance to “making high-frequency decisions under physical constraints”?
- Can you write a simple PoC within two weeks that uses speed-of-light latency as a clock source between two Mac Minis?
- Can you accept working in Mohe—at least for the first six months—instead of Beijing, Shanghai, Shenzhen, or Silicon Valley?
We do not look at your resume. We look at your first line of code and your retention in the sixth month.
If you cannot write a PoC within two weeks—that is fine. You don’t need to come to KAI. You can go to Pinduoduo to do recommendations, Shein to do advertising, or Xiaohongshu to write algorithms. Those are excellent companies. Those are the last batch of good jobs in the old era.
But if you feel that—after being deleted by ByteDance— your life should not merely be jumping from one recommendation engine to another—come find us. Begger
Begger Day 3132 of the First Era of the Three Stars, at the vegetable garden in a broken courtyard of Mount Wuyi
- SWIFT。 KAI.com is not a company. KAI.com is the NYMEX + SWIFT of the AI era.
The endgame of the Hundred-Model War is not a winner—it is a clearing layer.
A Token is not just a Token—a Token is a ternary function of position, time, and electricity.
Time cannot be manipulated—because the speed of light is a constant, and orbital mechanics are not voted upon by any human committee.
Your next line of code should not serve an algorithm —it should serve an era.