If you want the number without the preamble: in the United States in 2026, a machine learning engineer with roughly five years of experience earns somewhere between $220,000 and $550,000 in annualized total compensation — base salary plus the yearly value of equity plus target bonus. That band is wide on purpose. It is wide because at five years, your pay is set by two things that have nothing to do with your model-training skills: what level you were slotted into, and which tier of employer you work for.
Anyone who gives you a single national average for this role is selling you a number that describes nobody. The distribution outside big tech is bimodal — a cluster of well-funded AI companies and big-tech teams at the top, a much larger cluster of enterprise and mid-size employers well below them, and relatively little in between. A median sits in the empty middle.
The headline ranges by company tier
These are annualized figures: base + (equity value per year) + target bonus, for US roles at approximately five years of experience.
| Tier | Base | Equity / yr | Target bonus | Total comp | Confidence |
|---|---|---|---|---|---|
| Big-tech product ML (Google, Meta, Amazon, Apple, Microsoft) | $190K–$240K | $150K–$300K | $30K–$50K | $350K–$550K | Firm on level bands; equity varies with share price |
| Frontier labs (OpenAI, Anthropic, Google DeepMind research eng) | $250K–$400K | Often >2× base | Rare / folded in | $500K–$1M+ | Very low — see caveat below |
| Mid-size public tech | $165K–$200K | $30K–$80K | $15K–$30K | $220K–$310K | Estimate; most reports cluster here |
| AI startup (Series A–C) | $170K–$220K | Meaningful but illiquid | Usually none | $170K–$220K cash | Estimate; equity value unknowable |
A word on that second row, because it is the row everyone screenshots. Frontier-lab compensation is a tiny sample, it moves faster than any survey can track, and it is dominated by equity that is illiquid or only tradeable in periodic tender offers. Those packages are real, but they are not a benchmark, they are not a median, and you should not walk into a negotiation at a normal company citing them. If you do, you will read as uncalibrated.
How this squares with software engineer pay
Our software engineer salary page publishes level-by-level bands for the same five-years-of-experience population. Those bands are the reference point for everything below, and nothing in this post claims fresher or better data than that page:
| Company | Mid level (5 YOE typical) | Senior level (5 YOE possible) |
|---|---|---|
| L4 $270K–$380K | L5 $380K–$550K | |
| Meta | E4 $260K–$370K | E5 $370K–$570K |
| Amazon | SDE II $220K–$340K | SDE III $320K–$480K |
| Apple | ICT3 $245K–$350K | ICT4 $350K–$520K |
| Microsoft | 61–62 $210K–$300K | 63–64 $300K–$450K |
The site-wide implied big-tech band at five years is roughly $260K–$400K, stretching to about $550K for people who have already made senior. Notice that the ML band above ($350K–$550K) is not a separate, richer pay scale. It is the same ladder, weighted toward the upper half of it. ML engineers at these companies are disproportionately at the senior end of the mid/senior boundary and disproportionately in the highest-paying metros.
The correction: the "ML pays 30-38% more" claim is mostly a mix effect
This is the part worth reading twice, because it changes how you should plan a career move.
You will see the claim repeated everywhere that machine learning engineers out-earn software engineers by 30-38%. Aggregate title-level medians on levels.fyi and in survey data do show a gap of roughly that size. But an aggregate median is a weighted average of where people work, and ML engineers are not distributed like software engineers. They cluster at large tech companies, at AI-first startups with unusual funding, and in the Bay Area, Seattle and New York. Those employers and those metros pay more for every engineering role, including the backend engineer sitting next to the ML engineer.
Say it plainly: the ML median is higher partly because of where ML engineers work.
Hold the company and the level constant and the picture changes. Same employer, same level, the ML premium is generally in the 5-15% range, and it is narrowest at senior IC — some level-by-level comparisons put a Google L5 SWE and a Google L5 MLE only 4-7% apart. The premium that exists is real but modest, and it mostly reflects a tighter supply of people who can ship a production model rather than a structurally different pay band.
The practical consequence: switching from backend to ML inside your current company is unlikely to be a large raise. Switching from a mid-size enterprise to a big-tech ML team is a large raise — and most of that raise is the employer, not the title.
Base vs equity vs bonus, and the bonus nobody breaks out
Compensation pages usually publish base and equity and leave the bonus implicit inside the total. It is worth pulling out, because it is real money that people forget to count when comparing offers.
Take the Google L4 row above. Base plus equity is roughly $235K–$335K, while total comp is $270K–$380K. The residual is about $35K–$45K, which at those base levels is a target bonus of roughly 15%.
That is representative. At big tech, annual target bonuses generally run 10-20% of base. Two things follow:
- Amazon is the exception. Amazon does not pay a meaningful annual cash bonus. That value is folded into back-loaded RSUs (the notorious 5/15/40/40 vesting shape) plus year-1 and year-2 signing bonuses. The consequence is that an Amazon offer's first two years look competitive with peers and its steady state often does not. When you compare an Amazon offer to a Meta offer, compare year 3, not year 1.
- Bonus is a multiplier on base, not on total. A high-base, low-equity offer quietly earns more bonus than an equity-heavy one with the same headline number. This matters more at mid-size companies where equity is small.
At AI startups, an annual bonus is rare. Your cash is your base, full stop, and the equity is a separate decision.
Metro differences
Many companies moved to national bands after 2021; many others kept location tiers. For those that kept them, the spread between top-tier metros and tier-2 US cities is roughly 10-25% — meaningful, but smaller than the tier-to-tier spread between employers.
| Metro | Typical adjustment vs national ML band | Notes |
|---|---|---|
| SF Bay Area | Top of band | Highest concentration of frontier-lab and big-tech ML roles |
| New York City | Top of band, often at parity with Bay Area | Finance and adtech ML pull the top end up |
| Seattle | Slightly below Bay Area | Deep big-tech ML demand, no state income tax |
| Boston, Austin, Denver | 10–20% below top tier | Estimate; varies by company banding policy |
| Remote (national band) | Usually pegged to a mid-tier | Some employers still pay location-blind |
Treat every row here as an estimate. Banding policy is company-specific and changes without announcement.
What actually moves the number at five years, ranked
1. Level, not years
Five years of experience is the mid/senior ambiguity zone. The same résumé can be slotted as a Google L4 or an L5, a Meta E4 or an E5 — a gap the salary table above prices at roughly $110K per year. Being downlevelled is the single most expensive event in a five-year career, and it compounds, because the next employer will level you against your current level, not your original one. One downlevel can follow you for a decade.
Downlevelling is rarely a knowledge failure. It is usually a scoping failure — answering a system design question at component granularity when the interviewer was listening for ownership of the whole pipeline. I only noticed I was doing it after replaying a PhantomCodeAI mock round and counting how long it took me to mention monitoring and retraining. It was ninety seconds too late, every time.
2. Employer tier, once level is fixed
Look at the table again: Microsoft 61–62 tops out around $300K while Meta E4 reaches $370K, at nominally equivalent seniority. That is roughly $100K of spread between two big-tech employers at the same nominal level. Between big tech and a mid-size public company the spread is larger still. Almost nothing you do in your day job moves your compensation as much as this choice does.
3. Equity refreshers and the year-4 cliff
This is the most under-covered item at exactly this seniority, and it explains the confusing experience of getting a strong performance review and a smaller paycheck.
Your new-hire grant vests over four years. Annual refreshers are typically sized at 25-50% of the initial new-hire grant, and each refresher runs on its own four-year schedule. Do the arithmetic: by the time the original grant expires, the stack of accumulated refreshers frequently fails to replace it. Total compensation falls in year five with flat performance and no mistakes. Sequoia's work on equity refresh trends and Fearless Salary Negotiation's RSU refresher writeups both describe the same mechanic.
Two wrinkles worth knowing:
- Microsoft refreshers vest over five years, not four, which produces a quieter, longer dip across years two through four rather than one sharp cliff.
- Grants are denominated in dollars at grant date but paid in shares. The stock-price channel runs both ways. A rising share price can mask an under-sized refresher for years; a falling one turns a modest cliff into a large pay cut.
If you are at four to five years with your current employer, model your year-5 number before you decide you are happy. That modelling exercise is the reason a lot of people at exactly this seniority start interviewing.
4. Location
10-25% for the many companies that kept location-based bands, as covered above. Smaller than most people assume relative to the employer effect.
5. Negotiation, which means competing offers and essentially nothing else
The commonly cited 10-30% negotiation uplift is credible, but it is worth being precise about where it lands. It lands on sign-on bonus and the initial equity grant, which are one-time or four-year costs, not on base salary, which is band-constrained and sets future raises. A recruiter who cannot move your base by $5K can often move your sign-on by $40K.
And the uplift is driven by competing offers, not by rhetoric. The full mechanics — timing, scripts, what to say when asked for a number first — are in our salary negotiation guide. But at five years specifically, the highest-value negotiation happens earlier than most people realize: arguing your level, before any number is discussed. Once the recruiter has slotted you and made an offer, you are negotiating within a band. Before that, you are negotiating which band.
How confident should you be in each figure?
| Figure | Confidence |
|---|---|
| FAANG level bands and their level structure | Firm |
| Refresher mechanics and the year-4 cliff | Firm |
| Startup equity dilution of 50-70% across later rounds | Firm |
| Mid-size and startup cash bands | Most reports put them here; treat as estimate |
| 10-30% negotiation uplift | Widely reported, typically true, not guaranteed |
| ML-vs-SWE role delta at same company and level | Typically 5-15%; some sources narrower |
| Every ML band outside big tech | Range only — bimodal distribution |
| All frontier-lab numbers | Range only — tiny sample, fast-moving |
Sources worth checking directly rather than trusting a summary: levels.fyi (level-specific pages, and the title-level pages that produce the mix effect described above), Ravio's Compensation Trends 2026 and its software engineer salary trend work, Sequoia's 2025 equity refresh trends, and Fearless Salary Negotiation's RSU refresher overview.
Turning this into a plan
The compensation ladder rewards level, and level is decided in a loop. That is the uncomfortable link between a salary guide and interview prep: the $110K gap between E4 and E5 is settled by how you perform across four or five conversations, not by how many years you have accrued.
At five years the loop that decides your level is not the same loop a new grad sits. It is heavier on ML system design, on production trade-offs, on the "why did you choose that" follow-ups. If that is where you are, our machine learning engineer interview guide walks the loop structure, and the ML engineer question bank covers the archetypes that come up repeatedly.
When I was preparing for a senior-level loop, the thing that changed my outcome was not more LeetCode — it was discovering, in three mock rounds on PhantomCodeAI, that I consistently started designing before I had pinned down the evaluation metric. Interviewers read that as mid-level scoping. Fixing one habit was worth more than another fifty problems.
Two closing rules. First, never quote a point estimate — for yourself or to a recruiter. Quote a range and say what it is anchored to. Second, argue your level before you argue your number, because everything else in this post is downstream of that one decision.