Chapter contents · 13 sections
  1. 1. What ninety days actually finishes
  2. 2. Choose the scope, then start the clock
  3. 3. D1–D14: the Clarity Method audit
  4. 4. D8–D45: standards and context enter the work
  5. 5. From D15: let the live pilot happen early
  6. 6. Two kinds of pilot, one capacity gate
  7. 7. Settle leading indicators apart from operating results
  8. 8. What four companies show, and what they do not
  9. 9. bioby.ai: write the unverified parts as well
  10. 10. Accept the eight questions, item by item
  11. What to Do Monday Morning
  12. Closing: after the first round
  13. Chapter Acceptance Self-Check (against chapter acceptance standards)

Chapter 14Ninety Days: Completing the First Cycle of Gene Evolution

AI-native transformation is not the deployment of more tools. It is the shift of an organization from depending on a few people's experience, custom, and private information to depending on shared standards and decision-relevant context. What ninety calendar days can promise is finishing, inside a controlled scope, the first cycle that can be observed, reviewed, and iterated. It does not promise to complete the whole transformation, and it does not prove long-term ROI.

About 38 minContent date 2026-09-21

1. What ninety days actually finishes #

The previous chapter, on relationships, left a boundary: do not write changing another person as your long-term project. The same boundary holds when the field is the organization. If a principal-leader starts by requiring every department and every employee to change at once, the transformation soon becomes a campaign of persuasion. Meetings multiply. Slogans line up. The actual decisions, and the way work is done, stay as they were.

So the first round does not begin by changing the whole company. It begins with one real business unit, one class of decision that keeps recurring, or one workflow whose results can be observed. The organization first says clearly which judgments and facts it is already using, puts them into a piece of live work, and then watches how reality overturns, keeps, or revises them.

That is what completing the first round means. By day 90 the organization should leave at least six things:

  1. A Clarity Method audit that states the present problem, the baseline, the gaps in standards, the gaps in context, and the main uncertainty.
  2. An eight-layer standard v1 in which the clauses that bear directly on the pilot can already be executed, checked, and revised.
  3. Decision-relevant context with sources, versions, permissions, unknowns, and an owner for updates.
  4. At least one pilot that has entered a live business flow, and is not only a demo, a training, or a sandbox.
  5. A record of results and failures that separates a wrong standard, thin context, drifted execution, a mixed experimental design, and an external change.
  6. A revision of the standard and a plan for the next round, including owners, resource, stop conditions, and later review dates.

"Complete" means those six have closed a loop. It does not mean the pilot had to succeed. A pilot that stops on a line written in advance, that locates where the failure happened, that revises the standard and decides the next move, has still completed the first round. The reverse is also true. Strong short-term numbers without a baseline, a failure record, and a next edition of the standard mean the cycle never ran.

Why ninety days? Existing evidence cannot prove that it is the best cycle for every organization. Ninety days has one management job: to leave room for several rounds of feedback in high-frequency digital work, and to put an end to preparation with no deadline. From D1 to D90 there are twelve weeks of running plus six days of acceptance. When the business itself needs a longer clock, day 90 only checks the cycle and the leading evidence. Final results still settle on the real cycle.

Ninety calendar days for the first round: audit, standards and context, a live pilot, settlement, and the next round; three work lines run in parallel throughout

Figure: Live work begins around D15, before every standard is finished. Standards, context, and evidence are revised together while the work is running.

2. Choose the scope, then start the clock #

Dynamic lifecycle management locates the object first, then locates yourself, then matches the action. In a ninety-day manual the object is the decision or workflow you mean to change, and "yourself" is the organization's present stage, resource, capability, and risk. Skip that step and the roadmap becomes one table imposed on every company.

In a small organization the principal-leader owns the work directly, and the first round takes only one main pilot. A large organization should not try to cover the group on day one. It should take one division or business unit as the whole of this round, and let that unit's principal-leader own the result. A smaller scope does not make the goal lighter. Precisely because the range is one, standards, context, live results, and accountability can be made to line up.

A cash crisis, heavy regulation, and a period of major incident cannot copy the ordinary route. When the cash chain is deteriorating, first protect collections, critical delivery, and survival, and shrink the pilot to the shortest, reversible work that still moves cash. If the danger is still widening, do not start the ninety-day clock. In healthcare, finance, safety, employment, and other heavily regulated settings, professional owners should hold a hard stop: historical replay, a controlled environment, and dual review come before limited live traffic. When a major incident has already happened, first protect people, data, and customers. Completing the manual is not a reason to delay a stop-loss.

Accountability also has to be written on D1. At least five roles are needed:

  • The principal-leader or business-unit head holds final judgment. That person decides for whom what value is created, what ranks first when aims conflict, and which boundaries cannot be traded, and signs the coming into force, revision, and repeal of material standards.
  • A standard-domain owner maintains the standards in that domain. Maintenance is not solitary invention. The owner has to take in front-line facts, customer feedback, failure records, and counterexamples.
  • Front-line members supply the exceptions, patches, bad news, and unworkable clauses from actual work. They sit nearest to reality, and cannot be asked only to comply after a standard is published.
  • Evidence and risk owners separately maintain baselines, metrics, and source records, and the boundaries of data, compliance, security, and privacy. In a high-risk setting those jobs can sit with different specialists in the same organization.
  • An advisor may research, interrogate, chair a discussion, and draft. An advisor may not make the organization's final trade-off, hide disagreement on the ground, or sign a standard. Judgment cannot be transferred by a purchase contract.

This division of labor is not a democratic vote. Mission, ranking of value, and material risk still need a signature. A final signature is also not closed-door writing. If a principal-leader only writes a personal judgment, and then substitutes repeated speeches for front-line feedback, the organization still depends on that one person. Shared acceptance is this: relevant members can, in the principal-leader's absence, make an explainable decision from shared standards and necessary facts, and can push a revision when new evidence appears.

Stage card, D1–D3

  • Move: Choose one business unit and one real problem. Write the non-goals. Name the person with final accountability, the pilot owner, the standard-domain owner, the evidence owner, and the risk owner.
  • Owner: The principal-leader or business-unit head signs. A transformation coordinator organizes the work.
  • Deliverable: A ninety-day charter that states D1, D90, scope, roles, risk, permissions, and non-goals.
  • Leading evidence: The owner accepts the duty to stop. Necessary data can be obtained. The team has reserved capacity to execute and to review, not only time for meetings.
  • Stop if: There is no person with final accountability; necessary facts cannot be obtained; error is irreversible and there is no safe environment; the crisis is still widening.
  • Entry to the next stage: One observable real problem, one pilot owner, and a risk boundary written in advance.

3. D1–D14: the Clarity Method audit #

The audit does not ask how well people think they are using AI, and it does not count how many accounts have been bought. It looks for the places where the organization is still guessing: where there is no shared standard of what "good" is, where necessary facts sit in different versions, where the person accountable cannot obtain the information the decision needs, or where a result returns and the next round of work does not change.

Start with three kinds of live record—normal, failed, and exceptional. They can be recent customer deliveries, product decisions, hires, refunds, R&D releases, or resource requests. Normal samples alone let a lucky run be mistaken for a stable capability. Failures alone let every problem be explained as not enough effort. The three kinds together show which relationships keep repeating.

For each record, check at least six things:

  1. For whom what value was supposed to be created, and who held the final decision.
  2. What standard of judgment was used, and which parts were only one person's experience.
  3. Which facts the participants actually saw, and whether source, time, and version agreed.
  4. What was unknown, restricted, or in conflict, and whether permission matched accountability.
  5. On what caliber the result was accepted, and whether normal cases, failures, exceptions, and human patches were kept.
  6. After the result returned, which standard, which piece of context, or which permission changed.

Interviews can only supply leads. Where the budget actually went, where the principal-leader's calendar went, which version of material a decision record cited, how customers chose, and whether the same class of error repeated are the facts that correct those leads. Money, time, and judgment still belong to three layers of analysis: a company looks at money first, a person looks at time first, and an AI-native organization is finally limited by its capacity to produce reliable judgment. Consensus is not a new resource standing beside them. When people say consensus is missing, keep taking the claim apart: inconsistent standards, different context, unrecorded disagreement, a mismatch of permission, or feedback that cannot return.

A pilot is not chosen mechanically as "the three highest-frequency items," and the method does not require a median in every case. Selection has to look at value, frequency of reuse, the cost of error, observability, reversibility, the natural cycle, whether data can be obtained, and team capacity together. Statistical method likewise follows the question. Typical cases, total cost, volatility, extreme risk, and experimental difference may need different calibers. What matters is freezing the caliber before results appear, not locking the organization to one statistic.

The audit ends by choosing only the main uncertainty that most affects the next step. It may be "can this class of support request go to AI without a drop in quality," or "will the target customer keep paying for a result that is small but complete." An organization faces many unknowns at once. That is not a reason for the first round to try to extinguish all of them.

Stage card, D1–D14

  • Move: Re-examine three kinds of live record. Draw the present state of standards, context, process, permission, and feedback. Choose the main pilot.
  • Owner: The audit owner and the evidence owner organize. Front-line and customer interfaces supply counterexamples. The principal-leader confirms the limiting resource and the pilot.
  • Deliverable: A one-page status map, a baseline, a metric dictionary, a pilot hypothesis, guardrails, a stop line, a rollback, and review dates at 3, 6, and 9–12 months.
  • Leading evidence: Material facts have a source, a time, a version, and an owner. Unknowns still appear as unknowns. Unfavorable facts have not disappeared from the record.
  • Stop if: A same-caliber baseline cannot be obtained; the natural cycle is longer than ninety days and no credible leading indicator can be found; the scope is too large for an owner to be named.
  • Entry to the next stage: A frozen baseline, one main uncertainty, checkable metrics, and a stop condition written in advance.

The audit need not finish before the next step begins. From D8, high-value gaps already confirmed can enter a standard v0 and a context pack. Uncertain parts stay marked as hypotheses. That is not jumping the gun. It is letting the audit meet substitution and live work sooner.

4. D8–D45: standards and context enter the work #

A pilot standard that is to be handed to a person or to AI cannot stop at three sentences of "condition, action, acceptance." The minimum is still five fields: conditions of use, the action, evidence of acceptance, a stop condition, and a trigger for revision. Owner, version, and effective date have to be written as well. Otherwise the team knows how to go forward, and does not know when it must stop, or who revises when a new fact appears.

A context pack handles a different problem. It should at least list present facts, sources, versions, unknowns, restricted information, access permission, update frequency, and an owner. Unified context does not require the whole company to see everything, and it does not require people and AI to receive identical material. The caliber of acceptance is this: the people accountable for this decision obtain, within lawful permission, the same version of the necessary facts, and know what remains invisible. AI receives only the minimum permission the task needs.

From D8 to D21 the team first writes a standard v0 that bears directly on the pilot, then runs a substitution test on historical cases, a shadow run, or a low-risk sample. Find someone who did not take part in the drafting, but who holds the relevant duty, and ask that person to handle representative scenes from the standard and the context, independently. AI can do the same in a constrained environment. The test does not require identical answers. It records where they had to guess, when they knew they should stop, and what evidence proved completion. The points of guesswork are the gaps the next edition has to fill.

At the same time the eight-layer standard begins to form a v1:

  1. Mission states for whom what value is finally created, and what this round will not pursue.
  2. Belief states the causal hypotheses the organization is prepared to bet on for a long time, and what contrary evidence would force a second look.
  3. Behavior standards state the boundaries that cannot be traded: integrity, safety, legality, and a customer's informed choice.
  4. Thinking standards state how to separate fact, hypothesis, and preference, how to look for counterexamples, and how to check evidence.
  5. Judgment standards state how several feasible options are ranked, and under what conditions an option is abandoned.
  6. People-investment standards write a value hypothesis, a reasonable unit of acceptance, and a review window for a people investment. They do not force a shared result into a precise ROI for each person.
  7. Management standards assign, by concrete task, decision rights, information permission, a risk ceiling, and an escalation path. They do not paste a permanent rank onto a person.
  8. Learning standards state which results must be reviewed, where feedback lives, who publishes a revision, and how an old version exits.

v1 does not mean the eight layers are mature, and it does not require mission and a front-line checklist to be written at the same grain. The minimum is that none of the eight is empty, that the relations among them make sense, that the standards bearing directly on this round's pilot meet the five-field caliber, and that unverified parts are marked as hypotheses. When a layer cannot be written, the right move is to shrink the scope or to fill evidence, not to pad the table with slogans.

Getting a standard into the consensus cycle cannot rest on repeated speeches. Repetition can help memory. It cannot prove understanding, and silence cannot be recorded as agreement. Put favorable facts, unfavorable facts, limits of use, and known disagreements into one consensus record, and let the relevant members use it in real decisions. A person who objects need not recant. That person does have to know who currently decides, what action they own, and what new evidence can reopen the question.

Stage card, D8–D21

  • Move: Write a pilot standard v0 and a context pack v0. Build a permission matrix. Run historical, shadow, or low-risk substitution tests.
  • Owner: The standard-domain owner and the context owner draft. Front-line members supply exceptions. The principal-leader decides conflicts of value and the risk boundary.
  • Deliverable: Five-field standard cards, a context pack, an eight-layer skeleton, a decision record, and a template for version change.
  • Leading evidence: A person who did not draft can find the material, recognize unknowns, and name the stop condition. AI has not received surplus permission.
  • Stop if: The standard is still a slogan; a material conflict has no signature; sensitive information is overshared; an old version is still at a working entrance.
  • Entry to the next stage: v0 shows no uncontrolled material guesswork in substitution tests, and accountability and permission for the live pilot have been approved.

Stage card, D22–D45

  • Move: Use the first wave of live results to complete eight-layer v1. Record objections, unknowns, and version differences. Close old entrances.
  • Owner: The principal-leader signs. Each standard-domain owner maintains. Front-line, finance, data, and risk roles jointly supply facts.
  • Deliverable: Eight-layer standards v1, context pack v1, task-level decision rights and escalation paths, a revision record.
  • Leading evidence: Key members can make an explainable judgment in the principal-leader's absence, and can point to where a standard does not apply and where context is thin.
  • Stop if: The standard was written behind a closed door by the principal-leader; repeated speeches have replaced live use; objections and bad news have no safe entrance.
  • Entry to the next stage: v1 has been called in the live pilot, and has, on evidence from reality, already faced at least one judgment to keep, revise, or repeal.

5. From D15: let the live pilot happen early #

The largest problem in the old route was not a missing action. It was putting live verification last. Ten weeks of writing standards and building consensus, and three weeks of touching the business, give the organization at most one launch result. Failure, revision, and a second run are hard to live through. Those three weeks can prove that a team completed a project. They cannot prove that a cycle has been built.

The default route puts limited live business into the main pilot around D15. A heavily regulated setting can start with controlled live material, dual review, and assistive tasks that do not execute directly. It cannot postpone all verification in reality until D75. If a piece of work cannot produce any observable feedback inside ninety days, it is unfit to carry the first round's main pilot.

The first wave changes only what is enough to answer the main uncertainty. The team keeps the original baseline or another comparable caliber, and records quality, cycle time, errors, escalations, human patches, effects on customers or staff, and side effects that were not expected. Tool logins, generated word counts, and feature counts can help diagnose a problem of use. They cannot stand in for business evidence.

A stop line has to be written before launch. It may be harm to a customer, a safety incident, a compliance crossing, quality falling below an unacceptable range, damage to a stable core, a result that cannot be traced, or a run that can no longer answer the original question. Concrete thresholds are set in advance from the risk of the task and the baseline. There is no single number for every case. After a stop, do not only write "the pilot failed." Judge where the problem came from:

  • The standard was wrong: the original ranking of value, conditions of use, or caliber of acceptance does not hold.
  • Context was thin: facts were missing, a version was stale, permission was insufficient, or an unknown was not exposed.
  • Execution drifted: the standard and the context were usable, and the actual action did not happen as agreed.
  • The experimental design was mixed: too many conditions changed at once, and the result cannot answer the original question.
  • External conditions changed: the market, the regulation, the customer mix, or the underlying capability is already different.

Those five classes decide what to change next. If a failure cannot be returned to a standard, to context, or to experimental design, the first round has not increased the organization's judgment.

Stage card, D15–D35

  • Move: Put a limited live business flow into the main pilot. Record results, exceptions, patches, and unexpected consequences on the frozen caliber.
  • Owner: The pilot owner runs it. The principal-leader owns the decision to continue, pause, or roll back. The evidence owner keeps the source record.
  • Deliverable: A live-run log, an incident record, an interim result, and standards and context v0.1.
  • Leading evidence: There are already live cases and live exceptions. The stop line can actually fire. At least one result has changed a standard or a piece of context.
  • Stop if: A guardrail of safety, integrity, regulation, privacy, or the stable core has been crossed; quality has worsened and cannot be traced; the experiment can no longer answer the original question.
  • Entry to the next stage: At least one live business cycle has been completed, and the next wave's keep, adjust, roll back, or stop is explicit.

6. Two kinds of pilot, one capacity gate #

The field of the organization will call on both talent structure and product path. Every organization need not launch both rebuilds in the first round. The minimum deliverable is at least one live main pilot. The other kind starts only if owners, key people, data, and evidence chains can be kept apart, and then on a staggered clock. Otherwise it enters the next-round plan.

An execution-layer collapse pilot splits tasks from a role or a workflow. It does not begin from a headcount target. Item by item, the team judges which tasks execute on a clear instruction, which are handled by a mature standard, and which have to face exceptions, trade off value, and revise a standard. Whether AI or outsourcing can take a task is checked together on carrier, standard, context, acceptance, risk, and final accountability. After a migration, the people who remain receive Instruction, Guidance, Consultation, Authorization, or Delegation by concrete task. A person is not permanently marked S1, S2, or S3.

A product-path pilot locates demand first, then locates the organization's own capital, channels, technology, and delivery. The team freezes the main uncertainty that most affects the next decision, protects the stable core that is already creating value, and writes budget, term, guardrails, rollback, and stop. Answering one main uncertainty at a time does not mean a complex product can forever change only one variable. When factors interact, changing them one by one can produce a misleading conclusion. Staged release, a control, or a multivariable design can then be used, but the mix has to be admitted. Certain attribution cannot be invented (NIST/SEMATECH, 2012).

If a small organization puts the same core members on both pilots at once, it overdraws judgment, development, customers, and review capacity together. A large organization can have two owners and still lack independent attribution, if the two pilots compete for the same data, the same technical base, or the same customer result. The capacity gate asks only four things: whether the owners are different, whether there are enough key people, whether each has its own baseline and evidence chain, and whether one would change a main condition the other is testing. If any of the four cannot be answered, the second pilot does not start.

Stage card, D36–D60

  • Move: Stabilize the main pilot. Choose execution-layer or product path as the present main line. Assess whether to open the second kind of pilot.
  • Owner: The business-unit principal-leader makes the capacity decision. Each pilot already started has its own owner.
  • Deliverable: A mid-pilot review of the main line, a task-migration card or a product-trial card, and a written reason to open or not open the second pilot.
  • Leading evidence: Errors can be found, stopped, and recovered. The stable core has not been used to carry an unagreed experiment. The second pilot is not competing for the same bottleneck.
  • Stop if: There is no independent owner; key people are overloaded; the two pilots jointly change one result and cannot be told apart.
  • Entry to the next stage: Continue one main line, or, after the capacity gate is passed, run the second on a stagger. A path not opened enters the next round.

Stage card, D46–D75

  • Move: Run a second wave on the revised standards and context. Check whether the result can repeat after it leaves the drafter of the standard, the principal-leader's firefighting, and a special customer.
  • Owner: The pilot owner runs it. The evidence owner keeps observation, estimate, and inference separate for each pilot.
  • Deliverable: Standards and context v1.1, a second-wave record, and a keep / scale / roll back / stop decision.
  • Leading evidence: After substitution, judgment can still be completed. Rework, exceptions, and human patches begin to be explainable. Short-cycle results can be compared with the baseline on the same caliber.
  • Stop if: The gain appears only when the principal-leader personally fights fires; failure samples disappear from the record; the caliber of a metric is changed after the fact.
  • Entry to the next stage: After D75 the scope is no longer widened. Work turns to settling the evidence and designing the next round.

7. Settle leading indicators apart from operating results #

The easiest error in a ninety-day acceptance is to write what can be seen in time as final value. How often a standard is used, visits to a knowledge base, AI usage, training completed, and a member reciting a clause all show that an action occurred. Each still sits at a different length of causal chain from revenue, profit, customer results, and organizational judgment.

The first round should watch four kinds of leading evidence first:

  • Whether judgment can be traced: for a recent material decision, can the standard used, the version of the facts, the unknowns, the permissions, and the review date be found.
  • Whether shared material is usable: after substitution, can an explainable judgment still be made; where is guesswork still required; have old versions exited.
  • Whether live work has improved: on the same caliber, how quality, cycle time, errors, escalations, rework, customer feedback, and actual cost have changed.
  • Whether the cycle actually revised: whether a stop line was executed, and whether a failure changed a standard, a piece of context, a permission, or the next experiment.

Those pieces of evidence can support continue, adjust, roll back, or stop. They cannot automatically support "long-term ROI greater than 1." Three, six, and 9–12 months are the default management windows for staged review. They are not industry averages from a statistical study.

| Window of observation | What can be reviewed | What cannot be claimed from it | | --- | --- | --- | | D1–D90 | Whether standards and context entered the work; short-cycle quality, duration, error, exception, stop, revision, and some customer behavior | That the whole transformation is complete; that long-term revenue, profit, retention, culture, or ROI has been proved | | About 3 months | A first formal review of investment for an early organization or short-cycle work: value already realized, full cost, and key hypotheses | That every project should realize final value in three months; that short-term figures should be annualized by formula | | About 6 months | Whether reuse, retention, quality, unit economics, and adjustments of people and permission are stable across several business cycles | That one pilot has proved the method generally effective | | 9–12 months | Operating results across seasons, budget cycles, and complex collaboration; long-term customer results, risk, and organizational resilience | That concurrent change can all be attributed to this round of action |

From D61 to D84 the team stops widening the scope and sets baseline and result side by side. Facts that can be observed directly are listed alone. Model estimates and human inferences are listed separately. What is still unknown stays empty. A result that cannot be attributed reliably is kept as unknown. Proportions are not assigned in order to close the project.

Only from D85 to D90 does formal acceptance begin. A small organization should check every role that took a direct part. A large unit is checked in strata by role, risk, site, and type of task. Material risk, harm to a customer, privacy, compliance, and stop events have to be checked in full. Sampling is not a reason to skip them.

Stage card, D61–D84

  • Move: Freeze expansion. Compare baseline and result. Classify failures and alternative explanations. Publish the difference in the standard and later review dates.
  • Owner: The evidence owner and the standard-domain owner organize. Front-line, finance, data, and risk roles check. The principal-leader confirms the boundary of the conclusion.
  • Deliverable: A ledger of results and failures, a revision of the standard, unresolved hypotheses, a draft of the next round.
  • Leading evidence: Failures point to a concrete revision. Observation, estimate, inference, and unknown have been separated. Long-term value has not been booked in advance.
  • Stop if: The caliber of acceptance was changed after results appeared; a shared result was forced into a precise personal ROI; only success is reported, and stops are not.
  • Entry to the next stage: Material for the eight questions is in hand, and the next round's problem comes from the highest-value gap this round exposed.

Stage card, D85–D90

  • Move: Complete an evidence check of the eight questions, item by item. The principal-leader decides to continue, shrink, roll back, or stop. Freeze resource for the next round.
  • Owner: The principal-leader makes the final judgment. The evidence owner organizes the check. An independent interrogator may challenge, and may not sign in anyone's place.
  • Deliverable: An eight-question evidence pack, a decision that the first round is complete or incomplete, a next-round charter, and long-term review dates.
  • Leading evidence: Each question can point to actual material. A failed pilot has still left a next edition of judgment.
  • Stop if: Any minimum deliverable has no evidence. A total score cannot be used to declare the first round complete.
  • Entry to the next stage: The first round has ended. The second begins from a material gap in a standard or in context that has not yet been resolved.

8. What four companies show, and what they do not #

Whether ninety days is enough cannot be proved by four company stories. 与爱为舞 (Dance with Love), Transn, Klarna, and Air Canada sit in different industries, at different stages, and on different calibers of evidence. They did not take the same manual under a control. The material that can be checked presents four different situations: native design, rebuild of an existing organization, value compressed into cost, and a tool and a formal policy each using its own set of facts.

与爱为舞: human–machine collaboration designed from founding. In July 2025, in a public talk, founder Zhang Huaiting put the company's path as: first close a business loop and verify the scene, then let the model gradually assist or replace steps inside the loop. That path does not treat model capability as proof that the business already stands. It lets business results keep supplying data and constraint (Qiming Venture Partners, 2025).

Public material also describes the company putting product, R&D, operations, design, and sales into a human–machine way of working, and letting data from different steps flow back into one another. That material comes mainly from talks by company officers and from press. In November 2025, founder and COO Liu Wei disclosed that in about two and a half years from founding the company had completed four rounds of financing, totaling about $150 million, at a valuation near $1 billion; a technical lead had earlier disclosed monthly revenue in the tens of millions of yuan (36Kr, 2025). Financing, valuation, and revenue are all company or founder figures, unaudited, and they cannot prove that those results were caused by organizational design alone. 与爱为舞 shows that a native company can design business, data, and AI together from founding. It cannot prove that a native company needs no standards, and it cannot prove that other organizations can copy the path in ninety days.

Transn: an existing organization begins to change accountability and the entrance for feedback. Transn was founded in 2005 (Shanghai Stock Exchange, 2021). In March 2026, at a public event, the company created a CAIO and disclosed that more than twenty cross-department teams were running AI applications around hiring, coding, documents, customer service, and similar scenes (Hubei Daily, 2026). Later public material also describes an AI-native decision committee, a rule that a project must supply a runnable demo and state the business problem, and an "energy gold" mechanism accumulated from internal use and satisfaction (PingWest, 2026).

Those moves separately handle owners, business evidence, and feedback incentive. They sit closer to an organizational rebuild than buying a set of tools. A demo that runs is not proof that the right problem was solved. Usage counts and satisfaction cannot, on their own, stand for customer value. The available material comes mainly from company events and promotional reporting. It lacks before-and-after records of decision quality, customer results, revenue, cost, and failed projects. Transn therefore carries a sample of rebuild mechanism. It does not carry a proof of effect.

Klarna: after a cost metric pulled judgment off center, management corrected in public. Chapter 2 already recorded the sequence: the company first described success in volume handled and in cost, later said in public that cost had become too dominant an evaluation factor, service quality had fallen, and so, while keeping AI, it increased high-quality human support again. The meaning taken here for a first-round rebuild is only this—when the weights of evaluation are written wrong, the standard has to be revised by the result. In September 2025 the CEO added to Reuters that the company may have leaned too far toward cutting cost, and was shifting emphasis toward product and growth (Reuters, 2025). Whether the correction succeeds still depends on later customer and operating data.

Air Canada: one company gave customers two conflicting answers. Chapter 2 recorded Jake Moffatt's bereavement-fare case: the chatbot and the policy page it linked gave opposite rules; the company argued it should not be responsible for the bot; the tribunal held that all information on the website remained the company's (Moffatt v. Air Canada, 2024 BCCRT 149). The check taken here for ninety days is only this: the automated front door and the formal policy did not use the same version of decision-relevant context. "Context conflict" is a mechanism reading of the facts in the decision. It is not the tribunal's own phrase.

The four cases together support one limited judgment: after a tool enters the business, owners, value standards, relevant facts, and an entrance for feedback still have to be governed by the organization. They cannot estimate the average effect of a ninety-day route. That route is still missing one critical sample: a mid-size traditional firm that froze the caliber in advance, completed ninety days, and published process and operating results at D90, at six months, and at 9–12 months.

9. bioby.ai: write the unverified parts as well #

bioby.ai is this method's internal proving ground. What can be checked is a process that has not yet finished external verification: starting point, moves, failures, revisions, present results, and the parts still unverified.

Starting point. On the company's internal record, in early 2026 many material judgments still sat in the experience of the founder and a few members. Rework had already appeared in hiring, fundraising, brand design, R&D, and office outlay. Nearly twenty people were hired, one after another, into early commercial roles, and no suitable person remained. That result first shows a failure of hiring standard and of role design. It cannot be used to conclude that talent can only be screened.

Moves. Internal records show Clarity Method formed in February, Interference Method and four hiring standards in April, after which the new standards were first used to look again at sitting roles; from March to July the R&D unit redrew tasks, standards, and review accountability; in June, ROI discipline on people investment and some stop rules gradually became explicit. Public material has no month-by-month source ledger. Those dates are company self-report, not independently verified.

Failures and revisions. In the first month of expanding AI and junior members, quality in the R&D unit fell. Quality standards, task breakdown, and review accountability were then built. After the hiring standards formed, existing roles were checked first, not only later hires. Full construction of office meeting rooms was also paused because trade-off and a caliber of acceptance were missing. Those records show that new rules acted on sitting arrangements and on the principal-leader's decisions at the same time. A general ROI cannot be derived from a single internal event.

Present results. On the company's internal figures, the R&D unit had two skilled engineers in March 2026; four months later, one person mainly owned standards, breakdown, and review, nine interns took part with AI assistance, and the other engineer had left. Monthly people-and-tool cost rose from about 35,000 yuan to about 80,000 yuan; iteration and feature output were about five times the earlier level; response speed about three times; the production-environment bug rate fell by about 80%. The figures are unaudited. Full definitions of "output," "response," and "bug rate," and the source logs, have not been published. They cannot prove causation, and they cannot estimate what other organizations would obtain.

Still unverified. Stable use of the eight-layer standard in material decisions, the mix of time, applicability after the organization grows, and long-term customer and operating results all lack complete data. Time records for the core team are not yet finished. From the first item of method to some standards beginning to work together took the company about five months. That stretch has no control group. It cannot prove that other organizations can complete the same change in ninety days.

The case presents an internal process, company-reported results, and unverified parts. It does not carry a proof that the ninety-day method works.

10. Accept the eight questions, item by item #

D90 does not count how many tools were deployed, and it does not add the eight questions into a score. Each question has to point to a document, a version, a log, a live decision, or a running result. If any minimum deliverable has no evidence, the first round is incomplete. There is no passing by clearing six of eight.

Question 1: Has the Clarity Method audit been completed?

Evidence includes scope, a judgment of the limiting resource, samples of normal / failed / exceptional work, a baseline, a metric dictionary, a list of risks and permissions, and one main uncertainty. A mobilization meeting, a tool inventory, and spoken consensus cannot stand in.

Question 2: Is eight-layer standard v1 usable?

Each of the eight layers has a scope, an owner, a version, and hypotheses still to be verified. Standards that bear directly on the pilot have filled conditions of use, action, evidence of acceptance, a stop condition, and a trigger for revision, and have an effective date. Eight slogans, or a document that has not entered a working entrance, do not count.

Question 3: Do the relevant decision-makers obtain the same version of necessary context?

Evidence includes sources of fact, versions, unknowns, restricted items, permissions, an owner for updates, and a record of independent use on representative scenes. That everyone can see everything, that an email was sent, or that a knowledge base was built cannot pass on its own.

Question 4: Has at least one live pilot been run?

The pilot has a real customer, employee, business flow, or decision, and had, written in advance, a hypothesis, a baseline, a main uncertainty, guardrails, a stop line, and a rollback. A demo, a training, or a showing only in a report cannot satisfy on its own.

Question 5: Did final accountability and front-line fact both enter?

The principal-leader or unit head has signed material trade-offs. Front-line counterexamples, objections, bad news, and how they were handled also remain in the record. What an advisor did is likewise visible. A principal-leader writing behind a closed door, an all-hands vote that sheds accountability, or an advisor signing in anyone's place cannot pass.

Question 6: Were results and failures recorded as they were?

Same-caliber leading indicators, guardrail results, human patches, unexpected consequences, stops, and rollbacks are all there. Observation, estimate, inference, and unknown have been separated. Writing a ninety-day signal as long-term revenue, profit, or ROI cannot pass.

Question 7: Did the result change a standard, a piece of context, or the next action?

Evidence is a classification of failure, alternative explanations, a version difference, a reason for revision, an effective date, and the closing of an old entrance. Deciding, on evidence, not to change also counts. Writing only "strengthen communication" or "keep paying attention," with no field changed, does not close the cycle.

Question 8: Can the next round already be executed?

Unverified hypotheses, the next limiting resource, owners, resource, scope, a stop line, and review dates at 3, 6, and 9–12 months have all been written. A wish list with no commitment of resource is not a next-round plan.

When core roles are few, check in full the people who took a direct part in deciding, maintaining a standard, supplying context, and owning a pilot. When the business unit is large, sample in strata by role, risk, site, and type of task, and check live logs and decisions at the same time. Acceptance is not the principal-leader examining others and sitting out. The principal-leader's judgments, calendar, signatures, and records of correction sit in the evidence as well.

What to Do Monday Morning #

Monday does not begin with an all-hands transformation meeting. Do five things:

  1. Choose one real decision or workflow that recurs each week, whose result can be observed, and whose failure can be borne.
  2. Name one person with final accountability and one pilot owner, and write what this round will not do.
  3. Pull the most recent round of normal, failed, and exceptional records, and separate standard, context, process, permission, and feedback.
  4. Freeze the D1 baseline, the first main uncertainty, the leading indicators, and the first stop line.
  5. Put on the calendar a first live pilot around D15, a v1 review at D45, and eight-question acceptance at D90.

If the first step cannot find a real problem that can be owned, observed, and stopped, do not buy more tools. What the organization lacks at that moment is not capacity to execute. It is a judgment that can meet a test in reality.

Closing: after the first round #

The book began from execution and from retrieving knowledge getting cheaper. The question it has to answer at the end is not which job a person can hold forever unchanged. It is how a scarce capability in a person leaves that person and becomes an asset the organization can keep using and revising.

The last scarcity is the name of that question. On a person it is called adaptive insight: finding the gap in a standard, borrowing a verified standard, judging whether it can migrate, then correcting from the result. When a person writes the judgment down—conditions of use, evidence, boundaries, and a signal for revision—it begins to become a standard. When standards and decision-relevant context are used, questioned, and updated in common, personal adaptive insight begins to become organizational judgment.

At the end of ninety days there should not be a certificate that transformation is complete. On the table should be the first round's standards, context, results, failures, revisions, and the next-round plan. They do not guarantee that the next judgment will be right. They do keep the organization from having to start again on experience, seniority, and inertia alone.

The first cycle ends here. The next round begins with the real judgment you have to handle on Monday.

Chapter Acceptance Self-Check (against chapter acceptance standards) #

  1. Claim restatable in one sentence ✓: AI-native transformation is a shift from private experience to shared standards and decision-relevant context; ninety days finishes the first observable, reviewable, iterable cycle inside a controlled scope; it does not complete the whole transformation or prove long-term ROI.
  2. Ten-section spine and whiteboard figure ✓: what the first round finishes → scope and roles → Clarity Method audit → standards and context enter work → live pilot moved forward → two pilots and a capacity gate → leading evidence settled apart from operating results → four external cases → bioby.ai process self-report → eight-question evidence. Figure is D1–D90 with three parallel work lines; thirteen-week / last-three-weeks-only pilots are out.
  3. Evidence caliber ✓: 与爱为舞 from Qiming 2025 and 36Kr 2025, financing / valuation / revenue as company or founder figures, unaudited, not causal proof. Transn founded 2005 (SSE 2021); CAIO and twenty-plus teams from Hubei Daily 2026; committee / DEMO / energy gold from PingWest 2026; mechanism only. Klarna points back to Chapter 2 plus Reuters 2025; correction not booked as success. Air Canada from Moffatt v. Air Canada, 2024 BCCRT 149; "context conflict" as mechanism reading, not the tribunal's phrase. NIST/SEMATECH 2012 on interaction. bioby.ai as internal unaudited process; five months not proof others can finish in ninety days. Mid-size traditional longitudinal sample still missing.
  4. Case contract ✓: 与爱为舞 native design; Transn rebuild mechanism; Klarna wrong evaluation weights; Air Canada context conflict. MD Anderson, pass-six-of-eight, random three-person exam, personal ROI cards, Feynman-as-consensus-test, "with the map you walk faster," and a fifth form of last scarcity are out.
  5. Action menu ✓: Monday five steps; each stage card has a move, an owner, a deliverable, leading evidence, a stop, and an entrance; eight questions as eight kinds of evidence, not a score; principal-leader inside the evidence.
  6. Chapter boundary ✓: ninety days is a management time limit; second pilot only after a capacity gate; not cited as present proof; no new concept at the close; belief stays in the eight layers.
  7. Fluency ✓: rewritten in English voice from the Chinese authority; current prose-standard.