Attributing INP delays to third-party scripts: a vendor diagnosis method

A practical evidence chain for determining whether a tag manager, chat widget, analytics tool or experimentation script is contributing to a slow interaction.

A poor INP score tells you that users are experiencing slow interaction feedback. It does not identify the responsible vendor. A tag manager may be present without executing, a chat widget may trigger expensive first-party rendering, or a long task may have started before the user interacted at all.

That distinction matters when deciding whether to defer, reconfigure, replace or hold a third-party vendor accountable. Removing the wrong script can damage analytics, experimentation or customer support without addressing the underlying cause.

This article sets out an attribution method built around one evidence chain: affected interaction → INP phase → blocking work → script or dependency → vendor ownership → controlled change → field validation. The aim is not to prove that a vendor is the sole cause of every poor INP result. It is to establish whether a specific vendor configuration contributed to a materially slow interaction under defined conditions.

Start with the interaction, not the page score

INP measures the time from a user interaction until the next frame is presented, making it an interaction-level responsiveness metric rather than a simple page-load measure. The relevant unit of analysis is therefore a user action and its processing path: opening a filter, submitting a form, selecting a date or expanding a navigation menu.

The page-level INP value aggregates interaction observations for a page or origin. It helps identify that a problem exists, but it is too broad to attribute work to a particular vendor. Start by asking:

  • Which interaction target is slow?
  • On which route or template does it occur?
  • Which browser, device and network cohorts experience it?
  • What consent state, experiment assignment or logged-in state applies?
  • Does the delay occur consistently, or only under particular conditions?

In field data, capture the interaction target and, where governance permits, enough context to segment the result by route, browser, device class, consent state and experiment assignment. The web-vitals attribution guidance illustrates the type of interaction-level detail that can be collected. Availability depends on browser support, sampling and the way RUM instrumentation is implemented.

“The checkout page has poor INP” is a weak starting point. “Opening the delivery-date picker is slow for mobile users who have accepted analytics and personalisation cookies” is a testable diagnostic question.

Classify the INP phase before naming a vendor

INP can be examined as three timing components: input delay, processing duration and presentation delay. The Chrome INP breakdown documentation explains how these phases relate to the time before event processing, the event handler and related work, and the rendering of the next frame.

The phase changes the investigation:

  • Input delay: the interaction waited for the main thread. Look for work that was already running before the input, including a tag firing, analytics batch, chat update or first-party task.
  • Processing duration: the event handler and the work it triggers took time. Inspect the handler, callbacks, DOM updates and any vendor or first-party code called from that path.
  • Presentation delay: the browser took time to render the result. Investigate style recalculation, layout, paint and compositing, while checking whether a vendor callback caused the DOM or styles to change.

High input delay can result from unrelated main-thread work that was already running before the interaction. The code handling the interaction is not necessarily the source of the delay. The phase describes timing, not ownership or causation.

An investigation record might say: “Mobile filter-opening interaction; 180ms input delay; 45ms processing; 35ms presentation.” That is more useful than “third-party JavaScript is slowing the page”.

Use a trace to find the blocking work

Once the affected interaction and dominant phase are known, collect a browser performance trace that reproduces the same action. Chrome DevTools can help inspect interaction timing, event handlers, script evaluation, rendering work and third-party entities within a trace. Its Performance panel reference describes these capabilities.

Record the trace conditions explicitly:

  • browser and version;
  • device or CPU throttling;
  • network conditions and cache state;
  • route and page state;
  • consent choice;
  • experiment or personalisation assignment;
  • vendor configuration and release version, where known;
  • the precise interaction sequence.

Mark the interaction in the trace and follow the processing path backwards and forwards. For input delay, identify the task or tasks occupying the main thread immediately before the input. For processing duration, inspect the event handler and its descendants. For presentation delay, look at the rendering work after the handler.

A Long Task is a task occupying the user-interface thread for at least 50 milliseconds. It can delay input processing and other main-thread work, but the threshold does not make shorter tasks irrelevant. Several smaller tasks can also contribute to an interaction delay.

The Long Animation Frames API can expose frame-level timing and attribution, including source URLs, invokers, function names and execution timing, in supported environments. That can provide a useful bridge between an interaction and a script. The reported source may identify an entry point or event handler rather than the deepest library function responsible for the expensive work.

Do not treat a long task, or a vendor URL appearing inside it, as proof of causation. It shows that work occurred in the relevant time window. You still need to establish that the work was on the affected interaction path and that changing the suspected dependency changes the result.

Separate the five states of vendor evidence

Vendor attribution becomes clearer when evidence is recorded in stages rather than reduced to a yes-or-no conclusion. The following framework is a Plus IQ methodological proposal, not an established industry scoring system.

  1. Presence: the tag, script, iframe, worker or container is available on the page.
  2. Execution: the dependency actually ran before or during the affected interaction.
  3. Temporal overlap: its execution occurred in the relevant delay window.
  4. Mechanistic contribution: the trace links the work to the interaction through an event handler, callback, DOM change, rendering operation or blocking task.
  5. Demonstrated contribution: a controlled change reduces or moves the relevant work, and the improvement is later visible in representative field data.

These states prevent common diagnostic errors. A chat script can be present but inactive because the user has not opened the widget. An analytics library can execute during page load yet have no connection to a later filter interaction. Conversely, a tag manager may inject a vendor dependency that triggers a first-party callback. In that case, the tag manager is the delivery mechanism, not necessarily the owner of all recorded execution time.

Map ownership separately from execution. Record the tag-manager container, tag name, vendor domain, first-party wrapper, initiator and relevant dependency chain. A script loaded by a tag manager is not automatically “tag-manager CPU”. The expensive work may be in vendor code, a first-party callback or both.

Build the interaction attribution chain

For each suspected vendor, create one investigation record containing this chain:

  1. Affected interaction: target, route, cohort and observed field impact.
  2. INP phase: input delay, processing duration, presentation delay or a combination.
  3. Blocking work: task, event handler, script evaluation, callback, layout or paint operation.
  4. Dependency: URL, tag, container, iframe, worker, first-party wrapper or network-triggered callback.
  5. Ownership: the team or vendor responsible for changing the relevant code or configuration.
  6. Alternative explanations: other scripts, device mix, consent, experiment assignment, cache, network and first-party work.
  7. Controlled comparison: the intervention, matched conditions and result.
  8. Field validation: the cohorts and interactions monitored after implementation.

This record makes the conclusion auditable and exposes gaps. If the evidence jumps from “vendor present” to “remove vendor”, the diagnosis is incomplete.

Test the vendor without creating a false counterfactual

A controlled comparison is stronger than observational overlap. Depending on the suspected mechanism, compare:

  • vendor enabled versus disabled;
  • a feature enabled versus disabled within the same vendor;
  • normal execution versus delayed execution;
  • one trigger condition versus a narrower trigger;
  • the current version versus a matched previous version;
  • the current configuration versus a proposed configuration;
  • a vendor callback enabled versus a stubbed or no-op callback in a test environment.

The DevTools guidance on saving traces is useful when preserving comparable observations for review. Capture multiple runs where practical, keeping the interaction sequence and environment consistent.

A full disable is often useful for triage, but it can overstate the effect of a narrower remediation. Removing an entire chat widget may eliminate its script, iframe, network requests, event listeners and visual changes. That does not prove that delaying its unread-message poll, changing its trigger or disabling one feature will produce the same improvement.

Document what becomes non-equivalent. Disabling an experimentation script may change the page layout and therefore the interaction itself. Delaying analytics may reduce event volume or move processing to a later action. Removing a chat widget may change user behaviour. A test that improves INP while removing a business-critical function is not automatically a viable recommendation.

Also distinguish network and execution mechanisms. A slow vendor request may delay functionality without blocking the main thread, while a fast response can still deliver expensive synchronous work. A request can trigger later execution, so network timing may remain part of the causal chain rather than a separate explanation.

Use RUM and lab traces for different questions

RUM answers: Which users and interactions are affected in production? It can identify targets, routes, browser and device cohorts, consent states, experiment assignments and post-release changes when those dimensions are instrumented. It reflects real traffic, but sampling, privacy constraints and browser support limit the detail available.

A trace answers: What was the browser doing during this particular interaction? It can show tasks, event processing, script evaluation and rendering relationships. It supplies mechanistic detail, but it is a controlled observation of one environment and one path, not a representative estimate of every production user.

RUM alone generally cannot establish vendor causation because it may lack a complete runtime call stack, a counterfactual comparison and universal cross-origin attribution. A trace alone cannot establish production significance. The strongest conclusion comes when both point to the same interaction and mechanism:

  • RUM identifies a recurring slow interaction in a defined cohort.
  • A matched trace shows the suspected vendor or dependency executing on that interaction path.
  • A controlled change reduces or moves the relevant work under comparable conditions.
  • Post-release RUM shows improvement for the affected interaction and cohort.

Check the alternatives before assigning responsibility

Before escalating to a vendor, test plausible alternatives explicitly. The suspected script may be a convenient explanation rather than the correct one.

  • First-party work: a vendor callback may trigger state updates, DOM mutations or rendering in application code. The vendor may initiate the chain without consuming all of the recorded time.
  • Another dependency: tag managers, consent platforms, analytics libraries and experimentation tools can load or invoke one another.
  • Device and browser mix: a vendor may appear associated with slow sessions because it is more common in a mobile or older-browser cohort.
  • Consent state: consent can change whether a tag loads, when it loads and how it behaves. Record and match consent conditions. The Google Tag Manager consent documentation shows why consent signals affect tag execution, although implementations vary.
  • Interaction type: a click on a date picker may invoke a different path from typing into a search field or opening a chat window.
  • Network and cache: a delayed dependency can change the timing of later work without being the direct source of main-thread blocking.
  • Coincident releases: first-party deployments, browser changes, remote vendor configuration and traffic shifts can happen in the same period.

Vendor or container release timing is useful corroborating evidence. If a slow interaction worsened immediately after a vendor release, it helps narrow the investigation window. It cannot establish causation by itself because browser, application, consent and traffic variables may have changed at the same time.

A worked synthetic example

Consider an online learning platform where mobile users report that opening the course filter feels slow. RUM shows that the interaction has a high INP for users who have accepted analytics and personalisation cookies, but not for users who declined them.

A trace shows 140ms of input delay before the filter click is handled. A tag-manager analytics trigger runs immediately before the click, and its callback causes a first-party state update. The trace also shows a personalisation script evaluating during the same window.

At this stage, the evidence supports execution and temporal overlap for two candidates. It does not identify the sole cause. The team runs matched tests:

  • disable the personalisation feature while leaving analytics enabled;
  • keep both vendors enabled but delay the analytics callback until after the filter opens;
  • repeat both tests under the same consent state, device profile and cache condition.

If delaying the callback removes the blocking task but disabling personalisation has no material effect, the likely implementation action is to change the analytics trigger or callback timing. If the filter still performs poorly, first-party rendering or the personalisation script remains a live alternative. The defensible conclusion is that the tested analytics configuration contributed under the specified conditions, not that analytics is responsible for all poor INP on the platform.

Validate the remediation in two stages

After implementation, repeat the targeted trace first. Confirm that the suspected work has disappeared, moved outside the interaction window or been replaced by a cheaper path. Check that the interaction still behaves correctly and that the change has not simply moved the delay to a later click.

Then monitor representative RUM data. Compare the same interaction target, route and relevant cohorts before and after the change. Where possible, maintain the segmentation used during diagnosis: device, browser, consent state and experiment assignment. The field attribution data available through web-vitals can help keep the comparison focused on the affected interaction rather than only the aggregate page score.

There is no universal observation period or minimum sample size for every remediation. The appropriate period depends on traffic volume, effect size, seasonality and cohort coverage. A small site may need longer to observe a particular interaction; a high-volume route may produce useful evidence sooner. Record the decision rule before reviewing the result where possible, and separate genuine improvement from changes in traffic or browser mix.

What counts as sufficient evidence?

A practical attribution standard is reached when the evidence supports all or nearly all of these statements:

  • the affected interaction is identified rather than inferred from a page-level score;
  • the dominant INP phase is understood;
  • the suspected vendor or dependency executes on the relevant path;
  • the trace connects its work to blocking, processing or presentation delay;
  • ownership and first-party dependencies are mapped;
  • plausible alternatives have been tested or documented;
  • a controlled change alters the suspected work under comparable conditions;
  • field data confirms improvement for representative users.

Not every investigation will satisfy every point. Cross-origin iframes, workers, service workers, dynamically loaded resources and no-CORS scripts can limit detailed attribution. The Long Animation Frames documentation and Long Tasks specification describe relevant visibility limitations. In those cases, report the confidence level and the unresolved conditions rather than presenting a precise causal claim the evidence cannot support.

Third-party CPU totals and page-level performance reports are useful for triage. They can identify candidates for investigation, but they are insufficient to attribute a particular slow interaction to a vendor.

Conclusion: make attribution a chain of proof

The important distinction is between being present and demonstrating contribution. A script in the page, a vendor URL in a long task or a release that coincides with a regression can all be useful clues. None is decisive alone.

For SEO and engineering teams, the practical implication is to turn a performance concern into an implementation decision: identify the interaction, locate the blocking path, map ownership, test a proportionate intervention and validate the outcome in production. That produces a more useful recommendation than “reduce JavaScript” because it identifies what should change and how success will be measured.

Share this article

Found this useful? Pass it on.

Share on LinkedIn · Share on X