Menu

AI

Starting AI transformation inside the company: what Airbnb's case teaches other companies

Airbnb's CTO described putting AI into internal work first, then reshaping the customer experience. Checked against earnings material, six approaches other companies can apply.

Published
Boil time
9 min
Language
English · 日本語

TL;DR

  • Ahmad Al-Dahle, Airbnb's CTO, put AI into internal development first and builds customer-facing features on that machinery. The interviewer called this order "inside-out"
  • The numbers: about 60% of code written by AI, about 80% more features and improvements shipped year over year, and about 45% of support inquiries resolved by AI alone. The last two are also backed up in the earnings call
  • Six points ordinary companies can apply: make prototypes the deliverable, decide first which inquiries AI won't handle, keep what you learn from each project and reuse it, build evaluations per use case, hand first response to asynchronous agents, and keep people who can explain what AI produced
Contents
  1. Inside first, customers next: what the order means
  2. How far are the results backed up?
  3. From documents to prototypes as the deliverable
  4. Decide first which inquiries AI won’t handle
  5. An internal knowledge base that shortens development time for new services
  6. Choose models by evaluating them per use case
  7. Hand night-time response to asynchronous agents
  8. Keep people who can explain what AI wrote
  9. Caveats when reading

When someone says “become an AI-native company”, it’s hard to see where to start. Ahmad Al-Dahle, who led the release of Llama at Meta and became Airbnb’s CTO in January 2026, talks concretely about what that means in an interview with Latent Space. The interviewer calls this approach “inside-out”: use AI first for internal work to speed up development, then turn the same capability to the customer experience.

Airbnb has a market capitalization of about $93 billion, a different scale from an ordinary company. Even so, quite a lot of what he says can be copied regardless of headcount or budget. In this article I check the interview against earnings call records and other material, and focus on the points you can translate into your own work.

Inside first, customers next: what the order means

Al-Dahle says that at Meta he understood “the flow of raising model capability generation by generation”, and moved to Airbnb because he thought the next hard part would be “deploying models at scale”. His goals at Airbnb are two: “changing how people work” because of AI, and putting AI into production to “add value to the core experience”.

The specific order is as follows.

  1. Put AI into internal development processes and speed up how fast features ship
  2. Use the internal machinery built along the way to create new customer-facing services in a short time
  3. In customer-facing situations too, widen the range AI resolves in stages

If you reverse the order, you do your trial and error where customers can see it. Inside the company, failure has smaller impact and what you learn can be used on the next project. That is why an ordinary company too should make “internal work” its first target.

How far are the results backed up?

Checking the figures from the interview against the earnings call record (The Motley Fool’s transcript) gives the following.

  • About 80% more features and improvements shipped year over year: on the earnings call, CEO Brian Chesky said “the number of features and improvements we shipped was up about 80% compared with the same six months a year earlier”. This matches the interview
  • About 45% of support inquiries resolved by AI alone: Al-Dahle said “about half”, but the earnings call says “about 45% of inquiries that started with the AI assistant were resolved without a human agent”. According to Customer Experience Dive, it was 40% in the first quarter. The denominator is “inquiries that started with the AI assistant”, which you need to match when comparing with your own figures
  • Support cost: support cost per booking is down about 16% year over year. The earnings call attributes it in part to improvements in the AI assistant. It says “in part”, and doesn’t say all of it is AI’s doing
  • Time from concept to delivery: the earnings call says it was cut by up to 60% for some key initiatives
  • About 60% of code written by AI, and about 1.6 times as many pull requests handled per engineer: as far as I could confirm, these are not in the earnings materials and are written as Al-Dahle’s statements. The internal counting method isn’t shown either, so it is appropriate to read them as a rough guide

It seems better to read figures that can be checked from outside, such as ship counts and resolution rates, separately from figures that can only be verified by internal self-report. The same split works when you report the results of AI use at your own company.

From documents to prototypes as the deliverable

The first thing mentioned was a change in how the organization works. Previously, work moved forward by handing off to the next step: requirements, design in Figma, engineering implementation, validation in production. At Airbnb, he says, product, design and engineering teams changed to work while handling the prototype directly.

Al-Dahle says they moved from a state of over-producing “deliverables” such as requirements documents to one where “code and prototypes are the deliverables that inform decisions”. The explanation is that reducing the waiting time at each handoff was a major factor in saving time.

Applied to an ordinary company’s work, this can be read as follows.

  • Before writing a proposal, build something that works (a simple screen, or a prototype made by AI) and show it to stakeholders
  • Before circulating documents between departments, create a setting where people discuss while looking at the same screen or the same data
  • Keep documents as a record after a decision has been made

As the speed at which AI writes code goes up, the gap with the cost of writing documents narrows. There should be more situations where the order “polish the document, then build” ends up slower.

Decide first which inquiries AI won’t handle

Support was the first customer-facing area where AI was put in. Al-Dahle calls it the area where “deployment is hardest”, because mistakes have a big impact.

The approach has two features.

  • Test repeatedly on synthetic data before going to production: once a model and agent are built, you generate large amounts of synthetic data and try them out before they enter production
  • Deliberately decide which inquiries AI won’t handle: he says “while we are resolving about 50%, we are also conscious of the inquiries we’ve decided not to give to the agent yet”. The example given is inquiries involving safety

For an ordinary company’s help desk or internal inquiry desk too, it is safer to start by deciding “from where on a person takes over” rather than “how much to give to AI”. Putting a resolution-rate target first tends to pull toward making AI hold inquiries that should go to a person.

The earnings call also says the AI assistant is available in more than 50 languages. Voice (phone) support is also planned for later in 2026.

An internal knowledge base that shortens development time for new services

Airbnb’s grocery delivery and airport pickup are new services launched in early 2026. According to Al-Dahle, both are “similar kinds” of service, built on API integrations with partner companies, and an internal system for handling organizational context, “Everest”, supported their development. Everest builds and queries graphs using large language models, embeddings and AI search.

His explanation is that by accumulating what was learned from grocery in Everest, the airport pickup team cut the same work substantially. On the earnings call, Chesky also said “grocery took eight or nine months, and airport pickup we built in about six weeks”. However, the earnings call remarks don’t mention the link to Everest. The causal claim that it was “thanks to Everest” is appropriately read as Al-Dahle’s explanation.

As another effect of this system, he says “because there is a graph of context across the whole codebase, generalists can work even in highly specialized parts”.

An ordinary company doesn’t need to build something like Everest right away. Even in a small form like the following, the idea works.

  • Keep what you decided on the first project, where you got stuck and the partner’s specifications in a searchable form
  • On the second project, have AI read that before you start
  • Try it first in areas where similar projects keep coming up (integration with external services, routine requests and so on)

Effects are most likely in areas where similar projects keep coming, like this. Building a knowledge base for a one-off project may not pay for itself.

Choose models by evaluating them per use case

Airbnb is described as a “multi-model company”. It uses frontier models (each vendor’s most advanced models) together with open models, and does fine-tuning and reinforcement learning mainly in-house on open models. He says there are at least ten models customized for production use.

The selection criterion is the balance of cost, performance and latency (the Pareto frontier), and a different point is chosen for each use case.

Use caseTendency in the model chosenReason (Al-Dahle’s explanation)
CodingThe strongest frontier model availableLatency is tolerable, and the loss per defect is large
SearchA small, specialized modelUsage is at large scale and latency-sensitive

Evaluation is also per use case. For search, they measure with a set of search queries extracted from production, and for support, with a set of important production inquiries, including edge cases. He says that for narrow use cases they have managed to fine-tune a small model “to the point of exceeding frontier”. This is his claim, and no outside verification is shown.

You don’t need to do fine-tuning of open models yourself. Two points can be applied.

  • Decide not by “which model is smartest” but by “which model is good enough on samples of our own work”
  • Build the evaluation question set from records of real work

If you have evaluation samples on hand, you can decide whether to swap in each new model that comes out.

Hand night-time response to asynchronous agents

The next change Al-Dahle mentioned is asynchronous agents that run in containers and start from events. Airbnb has an internal conversational agent, “AirChat”, which handles organizational context through MCP. On top of that, many teams have begun automating on-call (first response to incidents).

The flow is as follows.

  1. A threshold in a monitoring tool (Grafana, for example) is crossed
  2. An agent starts and does the first-response triage
  3. If there is a fix, it submits it as a pull request. A human engineer reviews it
  4. If it judges the alert to be a false positive (flaky), the agent may also close the incident

He says that in the future he wants to extend it to monitoring fraud, trust violations, quality and software defects across the whole marketplace.

Here too, it is instructive that human review remains. An ordinary company can start in low-impact areas, such as first-pass triage of log anomalies or drafting routine incident reports.

Keep people who can explain what AI wrote

The last topic is the development of junior engineers. Al-Dahle’s concern is whether juniors can develop technical instincts and judgment while AI does much of the work, because senior engineers gained their judgment from running systems in production for years and failing along the way.

The measure he calls for strongly is that “even for an AI-generated pull request, the engineer can explain what was built”. If this premise is kept, the thinking goes, juniors can still learn the importance of interface design, architecture, and unit and integration tests.

This rule applies beyond software development. If you make it a condition that whoever submits AI-made materials, analyses or email text can explain “why it came out this way”, responsibility for checking doesn’t become vague.

He also says an organization that forms teams around goals and outcomes rather than around features is better suited to the AI era.

Caveats when reading

The interview is a company CTO explaining his own company’s efforts. The figures for results mix those backed up by the earnings call and those that exist only in internal statements. Note also that I couldn’t retrieve the original text of the earnings filings with the US Securities and Exchange Commission (SEC) in my environment, so I checked the earnings figures against the transcript and press coverage.

When applying this to your company, you don’t need to start all six at once. If you pick just one, it is realistic to start from these questions.

  • For the next project, can you build one prototype before the document?
  • Can you write out first the inquiries you won’t hand to AI?
  • Can you collect 20 evaluation samples from records of real work?

udon
Microsoft 365, ServiceNow, Copilot and more, tested first-hand and written up as practical notes with real bite.