Thirteen Years, Thirty-Three Reports, Millions In Public Money—and NYPD Is Still Underreporting Terry Stops

 

Executive Summary

More than thirteen years after federal remedial oversight began in the NYPD stop-and-frisk matters—Floyd v. City of New York, 959 F. Supp. 2d 540 (S.D.N.Y. 2013); Ligon v. City of New York, 925 F. Supp. 2d 478 (S.D.N.Y. 2013); and Davis v. City of New York, 902 F. Supp. 2d 405 (S.D.N.Y. 2012)—the Department is still failing to accurately document a substantial percentage of Terry stops. Nearly one-quarter of those encounters remain outside the formal reporting system, meaning thousands of constitutionally significant police encounters each year are not being captured in the manner necessary for meaningful supervisory, command-level, and judicial review.

That is a serious NYPD problem. It is also, by this point, a serious federal monitoring problem.

The purpose of a monitorship is not merely to observe an institution, describe its deficiencies, and perpetually document the same failures. A remedial structure is supposed to produce institutional change. It should identify systemic defects, force corrective action, create durable accountability, and ultimately make its own continued existence unnecessary. After more than a decade, however, the same basic problems remain: officers continue to misclassify investigative encounters, thousands of Terry stops remain undocumented, supervisors and command executives are not consistently held accountable, and the Department continues to operate without a sufficiently effective internal mechanism for ensuring compliance.

The numbers are difficult to dismiss. The estimated volume of undocumented Terry stops has declined over time, but thousands still disappear from the Department’s formal reporting system each year. The estimated compliance rate has improved from earlier years, but the most recent improvement is not substantial enough to establish that the Department has achieved stable, reliable compliance. Even accepting the progress reflected in the data, a system in which roughly one out of every four Terry stops may still go undocumented cannot reasonably be described as having solved the underlying problem.

The consequences extend beyond paperwork. When an encounter is not properly classified and documented as a Terry stop, it can evade the layers of review that are supposed to determine whether the officer had lawful grounds for the seizure. The supervisor may never review it as a stop. Command-level inspection may never capture it. Quality-control mechanisms may never flag it. The encounter may never become part of the dataset used to evaluate constitutional compliance. The classification decision therefore affects not only statistics but the ability of the Department, the court, and the public to understand what officers are actually doing on the street.

The deeper problem is that the institutional machinery necessary to address this deficiency already exists. The NYPD has policies governing investigative encounters, body-worn cameras, electronic evidence systems, reporting requirements, supervisory hierarchies, command-level review, internal auditing mechanisms, quality-assurance functions, and years of federal oversight. The Department has also been repeatedly informed that underreporting is occurring. The problem cannot credibly be described as lack of notice, lack of technology, lack of policy, or lack of data. The problem is accountability.

Officers who misclassify stops and fail to prepare the required documentation are rarely disciplined. Supervisors and command executives are not consistently held responsible for recurring failures within their commands. The Department is still developing more uniform systems to address conduct that should have been central to the remedial process from the beginning. That is not a minor administrative defect. It is the difference between a rule that exists on paper and a rule that actually governs behavior.

This is where the federal monitorship itself must be evaluated more critically. Audits can identify failure, but they cannot substitute for accountability. Statistical analysis can estimate how large a problem is, but it cannot make supervisors supervise. A corrective plan can assign another deadline, but it cannot by itself change organizational incentives. When the same problem has been identified repeatedly over many years and the remedial response remains another audit, another methodology, another plan, another deadline, and another round of monitoring, the public is entitled to ask whether the process has become better at measuring failure than correcting it.

There is also a structural problem in how compliance is measured. The monitoring system depends heavily upon the classifications assigned by NYPD officers to body-worn-camera recordings. When those classifications are inaccurate, the monitoring process must attempt to reconstruct what actually happened through sampling and retrospective review. The methodology has had to evolve to compensate for that problem, including the use of statistical reweighting to estimate Department-wide underreporting.

At the same time, the categories themselves continue to change. Videos classified as investigative encounters have declined sharply, while other classifications outside the traditional audit universe have increased. The monitoring process is now expanding again to examine categories that previously were not included in the audit. That does not establish deliberate manipulation, and it should not be presented as such. It does, however, expose a basic vulnerability in a compliance regime that depends upon the very institution being monitored to categorize the conduct that determines what gets monitored.

This is not an argument that the constitutional deficiencies identified in Floyd were imaginary, nor is it an argument that the NYPD should be left to police itself without meaningful oversight. The continued underreporting problem demonstrates why accountability remains necessary. The issue is whether the existing federal monitoring structure has produced results proportionate to its extraordinary duration and public cost.

A temporary remedy should have an endpoint. It should become less necessary over time because the institution being monitored has internalized the required standards and developed its own effective accountability mechanisms. When the opposite occurs—when persistent noncompliance continuously generates more monitoring, more analysis, more corrective plans, and more public expense—the remedial apparatus risks becoming self-perpetuating.

At some point, federal monitoring stops looking like a path toward institutional independence and begins looking like a permanent enterprise sustained by the very failures it has been unable to eliminate.

That is the question New Yorkers should now be asking.

I. Thirty-Three Reports Later, We Are Still Talking About The Same Problem

The number alone should give New Yorkers pause. In the September 28, 2026 Thirty-Third Report of the Independent Monitor: Report on Underreporting, the Monitor found that NYPD underreporting of Terry stops remains a persistent and significant problem more than thirteen years after the federal court imposed remedial oversight. The Department is still confronting one of the most basic issues at the center of that intervention: officers are conducting Terry stops that are not being accurately documented.

This is not an obscure technical requirement. A Terry stop represents a constitutionally significant seizure. An officer may not detain an individual merely because the officer wishes to ask additional questions or investigate a hunch. The encounter must be supported by reasonable suspicion, and once the interaction rises to that level, it must be recognized and documented as such. The reporting requirement is therefore inseparable from the constitutional standard. If an officer conducts what is functionally a Level 3 stop but records it as a lower-level encounter, the Department’s formal records no longer accurately describe what occurred.

That distortion has consequences. The Department’s data will understate the actual number of Terry stops. Supervisors reviewing formal stop activity may never see the encounter. Command-level inspections may miss it. Internal quality-control systems may not evaluate it. The data used to assess racial disparities, constitutional compliance, supervisory effectiveness, and command performance may all be incomplete. The failure to document a stop therefore does more than conceal a form; it can remove a constitutionally significant encounter from the very systems created to review it.

After thirteen years, however, the most troubling fact is not simply that underreporting exists. It is that everyone involved already knows it exists.

This is not the first time underreporting has been identified. The problem has been discussed repeatedly over the course of the monitorship, subjected to multiple rounds of auditing, examined through body-worn-camera review, and incorporated into successive compliance assessments. By now, underreporting cannot plausibly be characterized as a hidden weakness that federal oversight has only recently uncovered. It is a known institutional defect that has survived years of federal attention.

That history changes the question that should be asked.

During the early years of a monitorship, discovering a problem can itself constitute progress because the institution may previously have lacked the data, systems, or willingness to identify it. But the value of repeatedly identifying the same problem diminishes over time if the underlying conduct remains substantially unchanged. At some point, the measure of success must shift from whether the Monitor can find a deficiency to whether the remedial structure can cause the Department to correct it.

That distinction is especially important because institutional reform is easily confused with institutional activity. Government systems are capable of generating enormous quantities of work without producing an equivalent amount of change. Audits can be conducted, reports can be issued, recommendations can be made, policies can be revised, training can be delivered, committees can meet, statistical models can be refined, and corrective plans can be submitted. Every one of those activities may have some legitimate function. None, standing alone, demonstrates that the institution has actually changed.

The relevant outcome is behavioral. Officers should correctly recognize when an encounter becomes a Terry stop. They should document it consistently. Supervisors should detect misclassification. Commanding officers should know whether their personnel are complying. Persistent failures should produce meaningful consequences. The Department’s own internal controls should ultimately become reliable enough that outside monitoring is no longer necessary.

That is what a successful remedial process should produce.

Yet thousands of Terry stops are still estimated to go undocumented each year. The most recent compliance estimates still indicate that approximately one-quarter of stops may fall outside the required reporting system. The historical trend shows improvement from earlier years, but improvement alone cannot be the endpoint of a thirteen-year federal remedial process. The relevant question is whether the Department has reached the point where compliance is durable, internally enforced, and no longer dependent upon an external monitor to reconstruct failures through retrospective body-worn-camera review.

It plainly has not.

The apparent improvement in recent years also requires caution. Statistical movement from one year to another does not necessarily establish an actual institutional change when the differences fall within the range that could result from sampling variability. That matters because it would be easy to look at the upward trajectory and declare progress without confronting how much uncertainty remains in the underlying measurements. A monitoring regime of this duration should be able to demonstrate more than a favorable trend line. It should be able to demonstrate that the Department has fundamentally changed how it identifies, documents, supervises, and corrects Terry-stop activity.

The persistence of underreporting suggests otherwise.

The Department already has the infrastructure that should make accurate reporting achievable. Officers wear body-worn cameras. Encounters are categorized electronically. Stop reports are required. Supervisors are supposed to review officers’ activity. Command self-inspections exist. Quality-assurance mechanisms exist. Federal monitoring exists. The court remains involved. The City has spent years and millions of dollars building layer upon layer of review around the same conduct.

Yet the underlying problem remains.

That is why the Thirty-Third Report should not simply generate another discussion about whether NYPD officers are completing stop reports. The public already knows there is a compliance problem. The more important question now concerns the effectiveness of the remedy.

Federal monitoring was never supposed to become a permanent parallel bureaucracy whose success is measured by how accurately it describes continuing NYPD failure. The remedial process was supposed to change the institution sufficiently that extraordinary federal oversight would eventually become unnecessary. If the same core deficiencies remain after thirteen years, it is entirely appropriate to ask whether the monitoring model possesses the institutional leverage required to produce that outcome.

Otherwise, the logic becomes circular. The Department remains noncompliant, so the monitorship must continue. The monitorship continues, but substantial noncompliance remains. That continued noncompliance then becomes the justification for still more monitoring. Under that framework, there is no level of failure that can count against the effectiveness of the remedial structure because every failure automatically becomes an argument for extending it.

That may sustain a monitorship indefinitely, but it is not a meaningful measure of success.

After thirty-three reports, New Yorkers should be asking a more difficult question: how many times can a remedial system identify the same institutional problem before its inability to solve that problem becomes part of the problem itself?

II. The NYPD Problem Is Real—Which Makes The Monitorship Problem More Serious

There is no need to minimize the NYPD’s continuing failure in order to question whether federal monitoring has been effective. The two issues are not in tension. In fact, the persistence of underreporting makes scrutiny of the monitorship more important, not less.

Thousands of Terry stops appear to remain outside the Department’s formal reporting system each year. That matters because the initial classification of a police-citizen encounter determines what reporting and supervisory mechanisms follow. When an officer correctly identifies an encounter as a Level 3 Terry stop, the Department’s reporting system generally functions. The overwhelming majority of properly classified Level 3 stops have corresponding documentation. The much more significant failure occurs when an encounter that should be treated as a Terry stop is instead categorized as a lower-level interaction.

That finding is critical because it identifies the actual point of institutional failure.

The Department does not appear principally unable to process a stop report once the officer recognizes that a Level 3 stop occurred. The system breaks down earlier, at the classification stage. If the encounter is labeled as Level 2 rather than Level 3, the stop-reporting obligation may never be triggered and the interaction can pass through the Department’s systems without being treated as a Terry stop at all.

That is a far more serious problem than simple clerical noncompliance.

The distinction between a Level 2 inquiry and a Level 3 stop defines the boundary between an encounter in which an individual retains greater freedom to disengage and a seizure that requires reasonable suspicion. Misclassification therefore does not merely affect administrative statistics. It affects whether a constitutionally significant exercise of police authority enters the systems designed to assess whether that authority was lawfully exercised.

Once the encounter is misclassified, multiple safeguards may fail simultaneously. A supervisor reviewing stop activity may not see it because the stop was never formally recorded. Command self-inspection may not capture it because the encounter resides in the wrong category. Quality-assurance personnel may not review it because the data do not identify it as a stop. The federal monitoring process itself may only discover it later by sampling body-worn-camera footage and reconstructing what occurred.

That is not a sustainable model of internal accountability.

A mature police department should not require an external monitor to discover that an encounter was actually a Terry stop months or years after the fact by watching a sample of body-worn-camera videos. The Department’s own supervisory structure should identify the classification accurately when the encounter occurs or shortly thereafter. Supervisors should be reviewing officers’ conduct with sufficient rigor to detect obvious misclassification, and recurring problems within commands should generate managerial consequences.

The continued necessity of retrospective external reconstruction demonstrates how far the Department remains from that point.

There is also an important limitation in the existing monitoring process. The analysis necessarily begins with encounters that entered the body-worn-camera universe in the first place. If an officer failed to activate a camera during a Terry stop, that encounter may not be available for the same form of retrospective review. No one should speculate about how many such encounters exist without evidence, but the limitation matters because it means the monitoring system cannot claim to capture every undocumented stop. It can estimate what is missing among the recorded encounters available for examination; it cannot necessarily measure what never entered the system at all.

That limitation reinforces why the Department’s internal accountability structures matter more than any particular audit.

Federal monitoring can sample. It can estimate. It can compare classifications. It can identify patterns. It can expose weaknesses in supervision. But it cannot be present during every police encounter, and no external oversight system should be expected to function as a permanent substitute for competent internal management.

The goal of reform should therefore be an NYPD capable of identifying and correcting these problems itself. The Department should not need an external team to discover years later that officers were categorizing Level 3 stops as Level 2 encounters. Accurate classification should become part of ordinary supervisory practice, and command-level accountability should make persistent noncompliance institutionally costly.

That has not happened reliably enough.

The historical estimates demonstrate the seriousness of the problem. While the estimated volume of undocumented Terry stops has declined over recent years, the remaining numbers are still measured in the thousands. The direction is favorable, but the persistence is undeniable. Even the most recent estimate reflects a level of noncompliance that would be difficult to characterize as minor under any meaningful constitutional or managerial standard.

More importantly, those numbers arise roughly a decade into the federal remedial process.

That chronology matters. The same compliance rate means something very different in year one than it does in year thirteen. Institutional reform takes time, particularly in an organization as large as the NYPD, but time cannot become an unlimited excuse. At some point, the duration of the remedy must itself become part of the evaluation.

The public is therefore entitled to separate two questions that too often get collapsed into one. The first is whether the NYPD continues to have a Terry-stop reporting problem. It does. The second is whether the existing federal monitoring structure has proven sufficiently effective at correcting that problem to justify continuing the same model indefinitely.

The existence of the first problem does not answer the second.

Indeed, if persistent NYPD noncompliance is treated as automatic proof that the monitorship must continue unchanged, then the monitorship becomes insulated from any meaningful performance evaluation. The worse the Department performs, the stronger the argument becomes for continuing the same monitoring process, even if that monitoring process has been operating for years without producing durable compliance.

That is not a rational cost-benefit framework.

A remedial mechanism should be evaluated by what it changes, not merely by what it observes. Its existence may be justified initially by the seriousness of the underlying constitutional violations, but its continuation should ultimately depend upon whether it is producing measurable institutional results.

Thirteen years into federal oversight, the persistence of thousands of undocumented stops demands accountability from the NYPD. It also demands a serious accounting of whether the monitorship has become sufficiently effective to justify its duration, expense, and continued expansion.

The NYPD problem is real. That is precisely why the public should demand more than another sophisticated description of it.

III. The Real Failure Is Accountability

Beneath the statistical analysis, classification categories, sampling methodologies, and compliance percentages lies a much simpler explanation for why underreporting persists: the Department has not created meaningful consequences for the people responsible for it.

Officers who misclassify encounters or fail to prepare required stop reports are rarely disciplined. Supervisors and command executives are not consistently held accountable when those failures occur within their commands. More than a decade into the remedial process, the Department is still developing more uniform systems for auditing underreporting and imposing discipline for failures that have been known for years.

That is the central institutional problem.

The NYPD already has the legal standards. It already has reporting rules. It already has body-worn cameras, electronic evidence systems, supervisors, commanding officers, inspection processes, quality-assurance mechanisms, and outside oversight. The Department has also received years of notice that underreporting remains a problem. There is little reason to believe that another policy memorandum or another training presentation will solve conduct that persists because the institutional consequences for noncompliance remain weak.

Rules matter only when organizations enforce them.

A reporting requirement that officers can ignore or circumvent without meaningful consequence becomes less of a requirement in practice. A supervisory responsibility that carries no consequence when repeatedly neglected becomes administrative theater. A command structure in which executives are not held accountable for recurring patterns of noncompliance has little incentive to treat those patterns as operational failures requiring urgent correction.

That is why the distinction between auditing and accountability matters.

An audit is diagnostic. It tells an institution whether something is wrong, how frequently it may be happening, and where the problem appears concentrated. That function is valuable, particularly when a problem is newly discovered or its scope is unknown. But once the same deficiency has been identified repeatedly, the central question is no longer whether another audit can confirm it. The question is what happens to the people and commands responsible for allowing it to continue.

That is where this remedial process appears to have failed.

The Department has been told repeatedly that stop underreporting exists. The problem has been measured from different directions and through different methodologies. Body-worn-camera footage has allowed investigators to compare what officers recorded with what actually occurred. The disparity between correctly classified Level 3 stops and misclassified encounters makes clear where the system is breaking down. Yet the institutional response has not produced a reliable culture of accountability.

The federal monitoring apparatus must therefore confront the limits of its own approach.

A monitor cannot become the Police Commissioner and should not directly operate the Department’s disciplinary system. But after thirteen years, the relevant question is whether the remedial structure possesses enough leverage to require the Department to build a meaningful accountability system of its own. If the Monitor can repeatedly identify misconduct and supervisory failure but cannot produce an institutional environment in which those failures have real consequences, then additional reporting eventually produces diminishing returns.

That does not mean oversight should simply disappear. The continuing deficiencies provide substantial reason to reject any suggestion that the Department should be left entirely to its own devices. But maintaining ineffective oversight indefinitely is no more rational than abandoning oversight altogether.

The appropriate question is whether the remedy being used is capable of producing the required institutional result.

So far, the process has developed a familiar cycle. Officers misclassify encounters or fail to document them. The problem is detected through auditing. Another compliance deficiency is identified. The Department acknowledges the issue and develops another corrective plan. That plan is reviewed, another deadline is imposed, and additional audits are conducted. If the problem continues, the persistence of the problem becomes the justification for another round of monitoring.

There is nothing inherently improper about any individual step in that sequence. The problem emerges from the repetition. A remedial process can become self-perpetuating without anyone acting corruptly if every failure automatically produces another layer of the same process rather than a fundamental change in how accountability is imposed.

That is where the public interest becomes unavoidable.

A monitorship financed by taxpayers should not be judged simply by whether the Monitor performs work. Of course work is being performed. The relevant question is whether the work is producing durable institutional change commensurate with its duration and expense.

The public should not be satisfied by a system in which millions of dollars support increasingly sophisticated ways of proving that the same fundamental accountability failures remain unresolved. The point of oversight should be to force an institution toward compliance, not to create an indefinite professional apparatus for chronicling noncompliance.

The NYPD plainly bears responsibility for the underlying failures. Officers should accurately classify stops, supervisors should detect deficiencies, commanding officers should be held responsible for recurring problems within their commands, and executive leadership should establish consequences sufficient to make compliance the institutional norm rather than an aspirational objective.

But accountability cannot end there.

After thirteen years, the federal remedial process should also have to explain what measurable behavioral change it has produced, why the central accountability deficiency remains unresolved, what mechanisms have failed, how much longer the existing structure is expected to continue, and what objective conditions would finally make the monitorship unnecessary.

Without those answers, the federal oversight system risks becoming remarkably effective at one thing: documenting a problem that continues to sustain the need for federal oversight.

That may be a very good arrangement for the people paid to monitor the process. It is a far more difficult arrangement to justify to the public paying the bill.

IV. A Great Gig If You Can Get It: Failure Generates More Monitoring

There is a structural problem with any remedial system in which the persistence of failure becomes the justification for perpetuating the machinery assigned to correct it. That problem becomes particularly acute when the remedial apparatus is expensive, long-running, and funded entirely by the public.

The NYPD fails to comply. The failure is audited. The audit identifies deficiencies. Those deficiencies generate recommendations, corrective plans, deadlines, additional review, and another round of monitoring. If the deficiencies persist, that persistence becomes the basis for continuing the monitorship. The process can therefore sustain itself indefinitely unless someone eventually asks whether the remedial structure is producing enough institutional change to justify its continued existence.

That is the question that has been avoided for too long.

There is no need to accuse anyone of corruption or personal misconduct to recognize the obvious incentive problem. A monitorship of this size creates its own institutional ecosystem. Lawyers, consultants, statisticians, subject-matter experts, analysts, support personnel, and administrative infrastructure become attached to a remedial process that continues as long as the underlying institution remains noncompliant. The longer the problem persists, the longer the monitoring apparatus remains necessary. That does not prove anyone is deliberately prolonging the process, but it does create a structure in which failure itself generates more work for the people paid to oversee the failure.

That is precisely why the public should demand an objective measure of effectiveness.

The relevant inquiry cannot be whether the Monitor is busy. There is no doubt that substantial work is being performed. The relevant inquiry is whether that work is producing measurable institutional results proportionate to the amount of time and public money being consumed.

A monitorship should be judged by whether it is moving the institution toward independence from monitoring. The greater the success of the remedy, the less necessary the Monitor should become. That is the basic logic of any temporary corrective intervention. If, instead, the process continually expands because each new deficiency requires another audit, another methodology, another plan, and another period of review, then the remedial apparatus begins to function less like a bridge toward compliance and more like a permanent administrative enterprise.

That is where the phrase “public trough” becomes more than rhetoric.

Taxpayers are funding a structure that is supposed to correct institutional dysfunction. If the dysfunction persists year after year, the public should be entitled to ask whether the response is actually changing the institution or merely creating an increasingly sophisticated system for observing its failures.

There is a profound difference between accountability and dependency.

A successful monitor should make itself progressively less necessary by forcing the institution to internalize the standards, supervisory practices, and consequences required for compliance. An unsuccessful monitoring structure risks doing the opposite: it becomes the institution’s external compliance department, permanently responsible for detecting the failures that internal management should have learned to prevent.

That is not sustainable reform.

It is outsourcing accountability.

The danger becomes even greater when the monitoring process itself becomes the principal mechanism through which problems are identified. If the Department depends on an outside monitor to detect misclassification, underreporting, supervisory failure, and data anomalies, then the Department has not developed the internal capacity that the remedial process was supposed to create. The longer that dependency continues, the more difficult it becomes to claim that the remedy is working as intended.

A federal monitorship should therefore be subject to the same kind of performance scrutiny that would apply to any major public program. What was the problem? What was the intended outcome? What resources were committed? What measurable changes occurred? What deficiencies remain? What milestones were missed? What is the endpoint? At what point does continuing the same structure cease to be justified by the results?

Those questions become unavoidable after more than a decade.

It is not enough to say that the constitutional problems remain serious. That explains why accountability is still necessary; it does not automatically prove that the same monitoring structure remains the most effective or efficient means of producing it. If the NYPD continues to fail under the same remedial model, then the model itself must be open to reassessment.

Otherwise, there is no limiting principle.

The Department remains deficient, so monitoring continues. Monitoring continues, but the Department remains deficient. The continued deficiency then becomes the argument for continuing the monitoring. That cycle can repeat indefinitely unless someone finally asks whether the apparatus is producing enough change to justify its own continuation.

At some point, federal monitoring stops looking like a temporary remedy and starts looking like a business model.

That does not require a finding of bad faith. It requires only the recognition that institutions tend to perpetuate themselves unless they are forced to justify their continued existence. A monitorship should not be exempt from that basic principle simply because its mission is laudable.

The public should demand the same thing from federal oversight that federal oversight demands from the NYPD: measurable accountability.

V. The Monitor Is Still Discovering Problems With What It Is Monitoring

The weakness of the current structure becomes even clearer when the universe being monitored is itself unstable.

The monitoring system depends heavily upon how NYPD officers categorize body-worn-camera recordings. Those classifications determine which encounters fall into the audit universe, which ones are treated as investigative encounters, and which ones receive closer scrutiny. That means the reliability of the entire process depends in part upon the accuracy and consistency of classifications assigned by the very Department being monitored.

That is a significant vulnerability.

When an officer accurately classifies an encounter as a Terry stop, the reporting system generally works. The problem arises when the encounter is placed into a lower-level or otherwise different category. Once that happens, the stop may effectively disappear from the ordinary systems designed to identify and review it.

That is not merely a problem with officer judgment. It is a problem with the architecture of the monitoring process itself.

If the audit begins with Department-generated classifications, then any systematic shift in those classifications can alter what the Monitor sees. The system can only examine what enters the categories being sampled unless the Monitor continually expands the scope of review.

That is exactly what has happened.

The volume of recordings categorized as investigative encounters has declined sharply, while another classification outside the traditional audit universe has increased substantially. That shift may have innocent explanations, operational explanations, or something more concerning. The available information does not establish motive, and there is no basis to claim deliberate manipulation without evidence.

But motive is not the central issue.

The central issue is that the monitoring process had not been routinely auditing the growing category.

That should be troubling after this many years.

A mature remedial system should be capable of detecting when substantial amounts of activity migrate outside the categories that historically received scrutiny. If the monitoring structure must continually discover new classification patterns after they have already developed, then the process remains reactive rather than preventive.

That is a significant distinction.

Reactive oversight waits for a pattern to emerge, identifies it after the fact, adjusts the audit, and begins reviewing the newly relevant category. Preventive oversight is designed to detect and deter the institutional behavior before it becomes embedded.

Thirteen years into federal monitoring, the system should be much closer to the second model than the first.

Instead, the structure remains dependent upon retrospective reconstruction. A large category of body-worn-camera recordings changes over time, the monitoring process notices the shift, and the response is to expand future sampling.

Again, there is nothing inherently wrong with modifying an audit when new information emerges. Any responsible oversight system should adapt. The problem is what continual adaptation tells us about the durability of the underlying reform.

If the Department can repeatedly produce new classification patterns that require the Monitor to redesign the scope of review, then the problem is not simply whether the audit methodology is sophisticated enough. The problem is that the Department’s own internal controls are not sufficiently reliable to ensure that constitutionally significant encounters are being captured correctly at the front end.

That is the point the remedial system should have solved.

The Monitor should not have to function indefinitely as a forensic reconstruction unit, reviewing samples of body-worn-camera footage to determine whether the Department’s own classifications were accurate. That may have been necessary during the early years of oversight, when the scope of the problem was still being defined. It is much harder to justify as a permanent model of accountability.

The continued dependence on Department classifications also creates a deeper problem with measuring success.

If the categories change, then apparent improvements in compliance may partly reflect changes in what is being captured, categorized, or sampled. That does not mean the improvements are false. It means the monitoring structure must constantly account for the possibility that the data universe itself is changing.

That should make everyone cautious about declaring victory based on compliance percentages alone.

A reliable compliance system requires confidence that the underlying dataset is complete, consistently categorized, and resistant to the kind of classification drift that can alter what gets reviewed. Without that, sophisticated statistical analysis may produce increasingly precise estimates about an unstable universe.

That is not the same thing as institutional control.

The practical implication is straightforward. The problem is no longer merely whether officers complete stop reports. It is whether the entire information chain—from the street encounter, to body-worn-camera activation, to classification, to reporting, to supervisory review, to command oversight—operates reliably enough that constitutional compliance can be measured without an external monitor reconstructing the process after the fact.

The persistence of gaps in that chain demonstrates why the monitorship deserves scrutiny of its own.

Federal oversight was supposed to help create an NYPD that could accurately identify, record, supervise, and review its own conduct. If the Monitor must continually chase the Department’s changing classification practices to determine what should have been monitored in the first place, then the underlying institutional objective has not been achieved.

That should concern anyone interested in constitutional policing.

It should also concern anyone interested in whether taxpayers are receiving value for the enormous amount of time and money devoted to the remedial process.

VI. When The Monitoring Methodology Needs Another Methodology

There is nothing inherently wrong with stratified sampling, statistical weighting, or methodological refinement. Any serious audit of millions of body-worn-camera recordings must use sampling techniques, and any responsible analyst should adjust a methodology when the initial design does not accurately represent the underlying population.

The problem is not that the Monitor uses statistical tools.

The problem is what the continued need to redesign those tools tells us about the maturity of the monitoring process.

The original audit methodology intentionally oversampled certain categories of encounters because those categories were believed to present the greatest risk of underreporting. That approach made sense for detecting problematic encounters, but it also meant the sample could not be treated as representative of the Department’s full citywide stop activity.

A second layer of analysis was therefore necessary to reweight the sample and estimate what underreporting might look like across the entire population.

Again, that may be statistically appropriate.

But after years of auditing, the fact that the monitoring structure is still refining how to answer one of the most basic questions in the remedial process—how many Terry stops are actually going undocumented—should not be treated as insignificant.

It demonstrates that the federal oversight apparatus is still attempting to perfect the measurement system long after the underlying constitutional problem was identified.

That matters because sophisticated methodology can create an illusion of precision that obscures a much simpler institutional reality. Whether the estimated underreporting rate is twenty-five percent, twenty-seven percent, or some nearby number does not change the central fact that thousands of Terry stops remain outside the formal reporting system years after the Department was ordered to correct the problem.

At some point, the argument over methodology risks becoming another way to remain focused on measurement instead of accountability.

The larger limitation is even more fundamental. The audit can only examine encounters that exist within the body-worn-camera universe available for review. If an officer never activates the camera, that stop may never enter the dataset at all. There is no reliable method to estimate the full extent of conduct that was never recorded.

That limitation should not be exaggerated, but it should not be ignored either.

It means the monitoring system is necessarily working with an incomplete universe. The Monitor can identify misclassification among recorded encounters. It can estimate underreporting among those encounters. It can refine the statistical model. What it cannot do is reconstruct every encounter that never entered the system.

That exposes the inherent limits of external auditing.

No matter how sophisticated the statistical methodology becomes, it cannot substitute for an internal culture in which officers accurately classify encounters, activate cameras as required, prepare reports, and expect meaningful consequences when they do not. The further the system moves from those basic behavioral expectations, the more elaborate the external audit must become.

That is precisely the wrong direction for a mature remedial process.

Successful reform should simplify oversight over time because the underlying institution becomes more reliable. Data quality should improve. Classification errors should become less frequent. Supervisory review should catch problems earlier. External sampling should confirm compliance rather than continually uncover new categories of failure.

Instead, the current structure appears to be moving in the opposite direction. The methodology becomes more sophisticated because the underlying data remain unreliable. Additional categories are added because activity is shifting outside the original audit universe. Reweighting is necessary because the original sampling design cannot directly answer the citywide question. Future audits expand because new gaps emerge.

Every one of those adjustments may be defensible in isolation.

Taken together, however, they raise a broader concern: the monitoring process is becoming more complicated because the institution being monitored has not become sufficiently reliable.

That is a poor return on thirteen years of reform.

The public should not mistake methodological sophistication for remedial success. An increasingly elaborate system for estimating the size of a persistent problem is not the same thing as eliminating the problem.

There is an important distinction between knowing more and accomplishing more.

The federal monitoring process undoubtedly knows more today about NYPD stop reporting than it did at the beginning of the remedial period. It has more body-worn-camera data, more audit experience, more developed sampling techniques, and more sophisticated methods for estimating underreporting.

The question is whether all of that accumulated knowledge has translated into institutional behavior that makes the monitoring apparatus less necessary.

If the answer remains no, then the continual refinement of the methodology begins to look less like progress and more like evidence that the process has become trapped in an endless cycle of measurement.

A remedial system cannot justify thirteen years of operation simply by becoming better at counting the failures that remain.

At some point, the measure of success has to be whether the failures stop.

VII. Millions Later, What Exactly Did The Taxpayers Buy?

At some point, this stops being merely a discussion about constitutional compliance and becomes a basic question of public expenditure. Federal monitoring is not free. It is paid for by the people of New York City, who have funded years of professional services, legal work, statistical analysis, auditing, data review, technical consultation, administrative support, and repeated compliance assessments. That expenditure was not supposed to purchase an endless stream of reports describing the same institutional failures. It was supposed to purchase reform.

That distinction matters because government routinely measures expenditures by activity rather than results. Money is appropriated, professionals are retained, meetings occur, analyses are performed, recommendations are issued, and invoices are paid. Each step can be documented, and each can be described as evidence that something is being done. But none of that answers the question taxpayers are entitled to ask after thirteen years: what measurable institutional outcome did the expenditure actually produce?

The intended outcome was not difficult to define. NYPD officers were supposed to conduct investigative encounters within constitutional limits, accurately classify Terry stops, document them as required, subject those encounters to meaningful supervisory review, and operate within an internal accountability structure capable of detecting and correcting noncompliance. Over time, those reforms were supposed to become embedded sufficiently within the Department that extraordinary federal supervision would become unnecessary.

That is what taxpayers were supposedly buying.

Yet in 2026 the public is still being told that thousands of Terry stops may be undocumented, officers continue to misclassify encounters, supervisors and command executives are not consistently held accountable, the Department’s existing accountability mechanisms remain insufficient, the monitoring methodology itself has required further refinement, and additional categories of body-worn-camera footage must now be brought within the audit process. After thirteen years, the remedial system is still diagnosing failures that should have been addressed long ago.

That is where the concept of waste becomes unavoidable.

Waste does not require proof that no useful work was performed. A government program can produce competent work and still deliver poor value. A consultant can prepare a technically sound report, an auditor can conduct a legitimate review, and an expert can perform a statistically defensible analysis, yet the overall enterprise can still be wasteful if the expenditure fails to produce results proportionate to its cost and duration. The question is not whether the Monitor and the Monitor’s team worked. The question is whether New Yorkers received the institutional transformation for which they paid.

Thirteen years is long enough to make that comparison.

The public has financed a remedial structure that was supposed to make constitutional compliance more reliable. Instead, the system still requires an outside monitor to sample body-worn-camera footage, reconstruct encounters after the fact, determine whether officers classified them correctly, estimate how many stops went undocumented, adjust the statistical methodology to account for sampling limitations, and expand future audits when new classification patterns emerge. That is an enormous amount of remedial machinery devoted to determining whether the Department performed basic functions that its own supervisors and executives should already be capable of enforcing.

The persistence of that dependency should be treated as a cost, not merely as an operational inconvenience.

Every time the Department fails to develop sufficient internal capacity, taxpayers continue financing the external structure that compensates for that failure. Every time a new deficiency is identified, more professional services are required to study it. Every time the methodology must be modified, additional technical work follows. Every time a new category must be audited, the scope of monitoring expands. Every time another corrective plan is required, another implementation and review period begins. The financial consequence of institutional failure is therefore not abstract. Failure produces more publicly funded work.

That creates a perverse structure even if nobody is acting with improper intent.

The people paying for the monitorship do not benefit financially from its continuation. The City does not save money when the Department fails. The public receives no dividend when another year of monitoring becomes necessary. But the apparatus surrounding the monitorship continues to exist because the underlying failures continue to exist. The less successful the reform becomes at producing durable compliance, the longer the structure designed to oversee that reform remains necessary.

That is why the economics of the monitorship deserve the same scrutiny as its methodology.

No one should be permitted to answer the taxpayer question merely by pointing to the volume of work performed. The public is entitled to know how much has been spent, how those expenditures have been distributed, what measurable changes correspond to that spending, what deficiencies remain unresolved, what additional costs are projected, and what objective criteria will finally end the process.

Those questions are especially important because the cumulative public cost should not be treated as some incidental feature of a constitutional remedy. Public money is itself an accountability issue. Every dollar devoted to an ineffective or unnecessarily prolonged remedial mechanism is a dollar that cannot be spent on patrol staffing, training, technology, supervision, mental-health services, schools, housing, infrastructure, or any of the countless other legitimate needs competing for public resources.

A constitutional remedy is not exempt from cost-benefit scrutiny simply because its purpose is important. In fact, the importance of the purpose makes rigorous evaluation more necessary. If taxpayers are being asked to spend millions to produce sustainable reform, then the appropriate measure is whether sustainable reform has actually occurred.

The answer cannot simply be that more work remains to be done.

After thirteen years, that response becomes part of the problem.

The public deserves a complete accounting of what it has purchased. If the answer is thirty-three reports, repeated audits, revised methodologies, continuing underreporting, inadequate supervision, insufficient discipline, and another round of corrective planning, then New Yorkers are entitled to question whether the federal monitoring structure has delivered anything close to reasonable value for the money.

That is not hostility to constitutional policing. It is accountability.

If government can demand accountability from the NYPD, taxpayers can demand accountability from the people being paid to monitor it.

VIII. Another Plan, Another Forty-Five Days, Another Bill

The most revealing feature of the current remedial process may be how predictably institutional failure is converted into another procedural cycle. After years of monitoring and repeated findings concerning underreporting, the Department is once again expected to develop and finalize another plan for addressing the problem. That plan will then be reviewed, implemented, monitored, assessed, and almost certainly become the subject of future compliance analysis.

There is nothing inherently irrational about requiring a corrective plan when a problem is identified. Corrective planning is a normal part of institutional management. The difficulty lies in the history. When the same type of deficiency has survived more than a decade of federal supervision, another plan cannot be evaluated as though the remedial process began yesterday.

The question is not whether the next plan contains sensible provisions. The question is why the basic accountability structure that such a plan is supposed to create does not already exist.

The Department has known for years that Terry-stop underreporting is a serious compliance problem. It has known that officers can avoid the reporting system by misclassifying encounters. It has known that supervisory review becomes ineffective when the underlying encounter is not correctly identified. It has known that internal auditing and accountability mechanisms have not been sufficient. None of those issues emerged overnight.

Yet the institutional response remains another plan.

That is the bureaucratic cycle at the center of this entire problem. A deficiency is identified, a plan is demanded, the plan is reviewed, implementation follows, performance is measured, and when the deficiency persists, a new plan or additional corrective action becomes necessary. Each stage generates more monitoring activity and additional public expense, while the underlying behavioral problem continues to survive the process.

There comes a point when the procedural response becomes a substitute for actual reform.

Government institutions are especially susceptible to this because plans create the appearance of movement. A new protocol can be announced. A deadline can be imposed. A training requirement can be established. An audit can be scheduled. Progress can then be measured against implementation milestones rather than against the more difficult question of whether the underlying behavior has materially changed.

That distinction matters here.

The public should not care primarily whether the NYPD submits another plan within forty-five days. The public should care whether officers accurately classify Terry stops forty-five days later, whether supervisors actually detect misclassification, whether commanding officers are held responsible for recurring failures, and whether the disciplinary system imposes meaningful consequences when the same conduct continues.

Those are outcomes.

Everything else is process.

The danger of an extended monitorship is that process begins to become its own product. The production of another plan becomes evidence that reform is continuing. The review of that plan becomes another monitoring task. The implementation period becomes another period of federal oversight. Subsequent testing determines whether the plan worked, and if it did not work sufficiently, another modification becomes necessary.

That process can continue indefinitely.

The taxpayers financing it should be asking why.

There should be a point at which repeated planning without sufficient results triggers a different response. If the problem is that officers are rarely disciplined and supervisors are not being held responsible, the solution cannot simply remain another layer of reporting and another written plan. The remedy must create consequences strong enough to change organizational behavior.

Otherwise, each new corrective cycle becomes another public expenditure used to compensate for the same institutional weakness.

That is why “another forty-five days” carries more significance than it might appear to on paper. It represents yet another extension of a process that has already consumed more than a decade. The deadline may be administratively modest, but it sits inside a remedial structure whose duration is anything but modest.

At some point, deadlines should lead to endpoints.

The public should not accept a system in which every missed objective merely produces another date on the calendar. If a thirteen-year remedial process can continue simply by converting each unresolved deficiency into another implementation period, then the concept of temporary oversight becomes meaningless.

There must eventually be consequences for institutional failure beyond another round of monitoring.

Otherwise, the sequence becomes painfully predictable: another problem, another plan, another deadline, another audit, another report, and another bill.

That may be an excellent way to sustain a monitoring apparatus.

It is a terrible way to demonstrate that the remedy is working.

IX. Federal Oversight Was Supposed To Produce An Endpoint

Any extraordinary remedial structure should begin with an implicit premise: it will end.

Federal monitoring exists because ordinary institutional controls were deemed insufficient to produce compliance. That justification can support intrusive oversight for a period of time, particularly where constitutional violations are systemic and internal accountability has failed. But the legitimacy of a monitorship does not rest solely upon the seriousness of the original problem. Over time, it also depends upon whether the remedial structure is moving the institution toward the point where extraordinary intervention is no longer necessary.

That endpoint cannot remain theoretical.

A successful monitorship should progressively reduce its own importance. The institution should develop reliable internal systems, management should assume responsibility for functions previously driven by the Monitor, data should become sufficiently trustworthy that extraordinary auditing is unnecessary, supervisors should detect problems without outside reconstruction, and command executives should face consequences when systemic deficiencies emerge.

The monitorship should ultimately work itself out of a job.

If that is not happening, the public should know why.

Thirteen years is more than enough time to ask what the termination framework actually is. What specific conditions must the NYPD satisfy? Which deficiencies remain material enough to justify continued federal supervision? Which obligations have become institutionalized? Which have not? What measurable performance level would permit monitoring to contract? What performance level would permit it to end? How much additional time is reasonably anticipated?

Without clear answers, federal oversight risks becoming functionally permanent even if nobody formally declares it permanent.

That is particularly troubling because the continued existence of the underlying problem cannot, standing alone, justify indefinite continuation of the same remedy. Persistent noncompliance establishes that accountability remains necessary. It does not establish that the existing mechanism is the most effective means of producing that accountability.

That distinction has largely disappeared from the public discussion.

When a Department remains deficient, the reflexive response is that monitoring must therefore continue. But that logic assumes what it should be proving. If a remedial mechanism has operated for more than a decade and the same fundamental deficiencies remain, there should be serious consideration of whether the mechanism itself needs to change.

Perhaps monitoring should become narrower and more targeted. Perhaps command accountability should replace some forms of expensive retrospective auditing. Perhaps specific compliance functions should be internalized while independent verification is reduced. Perhaps sanctions or other consequences should attach to institutional failures that have survived repeated plans and recommendations. Those are policy questions requiring careful legal and factual analysis, but the important point is that continuation of the existing structure should not be treated as the only conceivable response.

There is a difference between oversight and inertia.

An oversight regime that continually reassesses its own necessity, narrows as compliance improves, and terminates functions that the institution can reliably perform internally is behaving like a remedy. An oversight regime that continually expands to compensate for new deficiencies, adds categories to its audits, revises methodologies, requires additional plans, and remains necessary because the underlying institution never fully assumes responsibility begins to resemble a permanent administrative structure.

That is not what court-ordered monitoring should become.

The absence of a meaningful endpoint also creates an accountability imbalance. NYPD officers, supervisors, and executives can be told that they must satisfy particular standards. The Department can be required to adopt policies, submit information, change procedures, and demonstrate compliance. But who is evaluating whether the monitoring structure itself has achieved its objectives efficiently? Who determines whether thirteen years is reasonable? Who asks whether the resources being consumed are proportionate to the results being achieved? Who decides whether another year of the same model is likely to accomplish something the previous thirteen did not?

Those questions should not be treated as attacks on judicial authority or constitutional reform. They are precisely the questions that should accompany any long-running use of public resources.

A monitorship should never acquire an entitlement to its own continuation.

The burden should increasingly shift toward demonstrating why continued oversight remains necessary in its existing form, what additional measurable outcomes are expected, how much those outcomes will cost, and when the public can reasonably expect the extraordinary structure to end.

Without that discipline, “temporary” oversight can become permanent simply because nobody requires it to justify the next extension.

That is how a remedy becomes an institution.

X. The Public Trough Needs Accountability Too

The central failure exposed after thirteen years of federal monitoring is not difficult to understand. The NYPD still has a serious accountability problem. Officers continue to conduct encounters that are not always classified correctly. Thousands of Terry stops may remain undocumented. Supervisory review does not consistently identify the problem. Command executives are not being held sufficiently responsible for recurring deficiencies. Internal controls remain inadequate enough that an outside monitor must continue reconstructing what happened by examining samples of body-worn-camera footage.

None of that should be minimized.

But neither should the performance of the federal monitoring apparatus be placed beyond criticism.

The public has spent years financing an extraordinary remedial structure designed to correct these very problems. It has paid for monitoring, auditing, statistical analysis, expert consultation, data review, legal work, corrective planning, and report after report. After all of that, New Yorkers are still being told that the Department lacks sufficient accountability, the auditing methodology requires continued refinement, new categories of recordings must be brought under review, thousands of stops remain undocumented, and another corrective plan is necessary.

At some point, the public is entitled to ask what it received for the money.

That question does not disappear because the work involves constitutional rights. If anything, constitutional reform should demand greater accountability because the stakes are higher. A remedial system that consumes millions in public resources should be able to demonstrate not merely that work was performed, but that the work produced measurable and durable change.

There is a profound difference between exposing dysfunction and curing it.

The federal monitorship has demonstrated that it can identify NYPD deficiencies. Thirty-three reports leave little doubt about that. The much harder question is whether it has developed a remedial model capable of eliminating those deficiencies and making itself unnecessary.

After thirteen years, that question cannot be avoided by commissioning another analysis.

Nor should the public accept the argument that continuing NYPD noncompliance automatically proves the need to maintain the same monitoring structure indefinitely. That is circular reasoning. If failure always justifies more of the same remedy, then the remedy can never fail. Every setback becomes evidence that it must continue, every unresolved problem becomes another reason for additional oversight, and every new deficiency creates more work for the apparatus assigned to address it.

That is precisely how a temporary intervention becomes a permanent seat at the public trough.

There is no need to allege corruption to recognize the structural problem. The monitorship exists because the Department remains deficient, and continued deficiency creates continued monitoring work. The professionals involved continue to be paid while taxpayers continue to finance the process. Whether anyone intends that result is beside the point. The arrangement should be judged by outcomes.

The NYPD should therefore face meaningful accountability for every persistent constitutional and reporting failure. Officers should accurately classify investigative encounters. Supervisors should detect and correct deficiencies. Commanding officers should be responsible for patterns within their commands. Executive leadership should be judged by whether those systems actually work.

But the same principle must apply to the remedial structure.

The public deserves a complete financial accounting of the monitorship, a clear explanation of the measurable institutional changes produced by that expenditure, an identification of what remains unfinished, a credible explanation for why those deficiencies remain after thirteen years, and a specific framework for ending federal monitoring.

What the public does not need is an indefinite subscription to reports explaining why the problem still exists.

Federal oversight was supposed to force institutional accountability. It should not become another institution insulated from accountability simply because its mission is reform.

There is a point at which another report is not progress, another plan is not reform, another methodology is not accountability, and another year of publicly funded monitoring is not evidence of success.

Thirty-three reports and thirteen years later, New Yorkers have every right to ask whether that point has already arrived.

If millions of dollars in federal monitoring still leave the City with thousands of undocumented Terry stops, weak supervisory accountability, evolving audit gaps, and another corrective plan, then the burden should no longer be on taxpayers to unquestioningly finance the next round.

The burden should be on the monitorship to demonstrate why they should.

About the Author

Eric Sanders is the founder and president of The Sanders Firm, P.C., a New York-based law firm focused on civil rights, immigration, employment discrimination, police misconduct, and other high-stakes matters. A retired New York City Police Department (“NYPD”) officer, he brings a rare inside perspective to the intersection of government power, public institutions, enforcement discretion, and constitutional accountability.

Over more than twenty years, Eric has counseled thousands of clients and handled complex matters involving police use of force, sexual harassment, retaliation, systemic discrimination, immigration consequences, and related civil-rights violations. His immigration practice focuses on family petitions, green cards, citizenship, removal defense, humanitarian protection, waivers, appeals, and complex status issues. He graduated with high honors from Adelphi University and earned his Juris Doctor from St. John’s University School of Law. He is licensed to practice in New York State and in the United States District Courts for the Eastern, Northern, and Southern Districts of New York.

Eric has received the You Can Go to College Committee Foundation Humanitarian Award, The Culvert Chronicles 2016 Man of the Year Award, the National Association for the Advancement of Colored People (“NAACP”)—New York Branch Dr. Benjamin L. Hooks “Keeper of the Flame” Award, and the St. John’s University School of Law Black Law Students Association (“BLSA”) Alumni Service Award. He is widely recognized as a leading New York civil-rights attorney and a prominent voice on evidence-based policing, institutional accountability, equal justice, and rights-based immigration advocacy.