What Is E/M Leveling and How AI Gets It Right

Nearly 42% of physicians reported at least one symptom of burnout in 2025, according to the American Medical Association’s national physician comparison report. Documentation and billing remain named drivers of that number every year the survey runs. E/M leveling sits at the center of the problem because it is the single coding decision that determines what a visit actually pays.

E/M leveling is the process of matching a patient encounter to the correct Evaluation and Management code, from 99202 through 99215, based on medical decision-making or total time. Get the level wrong, and a practice either loses revenue on every undercoded visit or invites an audit on every overcoded one. Notiro built its coding engine around this exact decision point because ambient scribing alone never solved it.

What E/M Leveling Actually Means

Since 2021, E/M leveling for office and outpatient visits has run on two tracks. A physician can select a code level based on Medical Decision Making, which weighs the number of problems addressed, the data reviewed, and the risk of complications. Or the physician can select based on the total time spent on the encounter date.

CMS reaffirmed this structure for 2026, and only two of the three MDM elements need to be met for a level to apply. The core framework has not changed since 2021, but CMS clarified several interpretation points in its September 2025 Evaluation and Management Services guide, including how a stable chronic illness should be documented before it can support a higher level.

CPT itself keeps expanding around this framework. The AMA’s CPT 2026 code set added 288 new codes and 418 total changes, several of them tied to AI-assisted and remote monitoring services that now intersect with E/M clinical documentation. A physician coding manually is expected to track it all while running a full patient schedule.

Why E/M Leveling Breaks Down in a Real Practice

The theory is clean. The reality, at the end of a twelve-patient day, is not.

One documented case from Medical Economics found a spine specialty practice billing 94% of its visits under code 99213, the lower of two common established-patient levels, despite a benchmark that supported far greater use of 99214. The gap came to nearly $8,000 a month in underpayment from one clinic on one code pair. That is not fraud. It is a physician choosing the safer, lower code because auditing a chart takes less time than defending it.

This is the pattern behind undercoding across the industry: not intent, but time pressure colliding with a genuinely complex rule set. ICD-10 alone carries more than 70,000 codes, and CPT adds another 10,000, according to CMS. Manually cross-referencing that volume against a rushed visit note is a structural setup for a lower-than-earned code, every time the physician defaults to caution.

What AI Evaluation and Management Coding Changes

AI evaluation and management coding does not remove the MDM-versus-time decision. It applies that decision to the actual content of the visit, at the moment the note is written, instead of hours later from memory.

An AI-powered E/M coding system reads the ambient transcript and the finished note, then maps the problems addressed, the data ordered or reviewed, and the risk of the treatment plan against the current MDM table. Because the analysis is conducted against the actual conversation rather than a recalled summary, it captures the complexity a rushed physician would otherwise leave out of the chart. That is the direct fix for the 99213-versus-99214 pattern documented above.

This is where E/M documentation AI and coding automation diverge from a generic ambient scribe. A scribe who only writes the note still leaves the physician to level it manually after the visit ends.

How Automated E/M Level Selection Works Inside Notiro

Notiro’s ICD-10 and CPT coding automation generates suggested diagnosis and procedure codes directly from the visit audio and the note it produces, then presents them to the physician for acceptance or adjustment before the one-click sync to the EHR. Automated E/M level selection follows the same logic: the system proposes a level supported by what was actually discussed and decided, and the physician confirms it.

No Tier 1 ambient scribe on the market, including Freed AI, Heidi Health, and Nabla, currently ships this kind of coding automation. DeepScribe offers something comparable but charges $350 to $500 per provider per month, positioning it for enterprise health systems rather than solo or small-group practices that do most coding by hand today.

Physicians who adopt AI scribing tools earn approximately $3,000 more per year and see about one additional patient per week, according to UCSF research cited in industry coding materials. That gain comes from two sources working together: faster documentation and code that actually reflects the visit.

AI Coding Accuracy: What “Getting It Right” Actually Requires

An honest answer matters more than a clean headline here. CMS guidance for 2026 is explicit that an AI coding assistant for physicians cannot determine medical necessity on its own, and that providers remain responsible for reviewing, verifying, and authenticating every code before it leaves the chart.

That is not a limitation. Notiro hides behind a marketing claim. It is the design. AI coding accuracy comes from pairing algorithmic pattern recognition, which captures the complexity a tired physician might miss, with a physician’s clinical judgment, which catches anything the algorithm misses. Neither one, alone, gets a multi-problem, time-pressured visit coded correctly.

The physicians most skeptical of AI medical billing software are usually right to ask how it performs on the hard cases: the diabetic patient with a new GI complaint and a medication change in the same fifteen-minute slot. That is precisely the visit where automated leveling earns its place by surfacing every problem addressed, rather than the one the physician remembers writing about.

Choosing AI Clinical Documentation and Coding Built for the Full Visit

A practice manager evaluating tools should ask one direct question: Does this system code the visit, or only record it? Losing one physician to burnout costs a practice $500,000 to $1,000,000 in recruitment and lost productivity, per AMA analysis, and undercoding quietly compounds that loss every month a practice stays understaffed.

Notiro was built around the full clinical day: patient intake before the visit, ambient scribing during the visit, and ICD-10/CPT coding after the visit closes. For the solo physician trying to reclaim an evening, the small-group practice manager tracking E/M distribution against specialty benchmarks, and the nurse practitioner whose documentation burden mirrors a physician’s without the tools built for one, that full sequence is the difference between a scribe and a solution.

The Coding Decision Physicians Shouldn’t Have to Guess At

Undercoding will keep costing practices money for as long as leveling depends on a rushed physician’s memory of a twelve-patient day. AI evaluation and management coding does not remove the physician from that decision. It removes the guesswork by proposing the level the visit actually supports before the note is even closed.

Notiro is the only AI medical scribe built to carry that decision through the entire encounter, from the first word of the visit to the code that reaches the EHR. Every other Tier 1 scribe stops at the note. Notiro codes the visit.

Start Your Free Trial

Undercoding quietly drains revenue from every rushed visit, billing it at a safer, lower level. Notiro’s automated E/M level selection reads the encounter as it happens and proposes the code it actually supports, so the physician reviews the level instead of guessing. Start a free trial at notiro.ai, no IT setup and no enterprise contract required.