ooligo
ENTRY TYPE · definition

eDiscovery

By Marius Bughiu Last updated 2026-08-10 Legal Ops

eDiscovery (electronic discovery) is the process of identifying, preserving, collecting, processing, reviewing, and producing electronically stored information (ESI) in response to litigation, regulatory investigations, or internal investigations. It is the legal industry’s largest data-engineering problem: a single second-request response can ingest tens of millions of documents, and one missed privileged document can become a waiver argument.

eDiscovery is not legal research, and it is not contract review. Research tools search published law; eDiscovery searches your own company’s mailboxes, Slack channels, and file shares. Contract tools like Ironclad read documents you are about to sign; eDiscovery reads documents you already sent. It is also not a synonym for legal hold — the hold is one stage inside eDiscovery, not the whole process.

The EDRM model

The Electronic Discovery Reference Model (EDRM) is the industry’s map of the workflow. The nine-stage version is the one currently in operational use and the one to cite in an ESI protocol today:

StageWhat happens
1. Information governancePre-litigation: data retention policies, hold readiness
2. IdentificationWhat custodians and data sources are in scope?
3. PreservationIssue legal holds; freeze deletion
4. CollectionPull data from email, file shares, Slack, mobile, cloud apps
5. ProcessingDeduplicate, extract metadata, OCR, normalize formats
6. ReviewAttorneys, or AI, tag documents as responsive, privileged, hot
7. AnalysisBuild the case narrative from reviewed documents
8. ProductionProduce responsive, non-privileged docs to the requesting party
9. PresentationUse produced documents at deposition or trial

EDRM published a draft EDRM 2.0 for public comment on 30 June 2026, built by roughly 150 practitioners. The comment window closed 30 July 2026 and trustees are working through submissions before final publication, so the table above still governs — but three changes are already visible and worth planning against.

  • Information Governance stops being stage 1 and becomes the foundation layer beneath the whole lifecycle. The argument is that governance decisions made years before a complaint determine every downstream cost.
  • Identification, Preservation, Collection, and Processing are grouped as Data Acquisition, because they now run concurrently rather than in sequence.
  • Disposition joins as a discrete final phase — what happens to matter data after the case closes, which the nine-stage model never named.

Analysis is also redrawn as continuous across every phase instead of a single step wedged between review and production.

Most eDiscovery platforms — Relativity, Everlaw, DISCO, Logikcull — cover collection through production. The stages at either end are handled by Microsoft Purview, Google Vault, or dedicated retention tooling, which is exactly the gap EDRM 2.0 is pulling into frame.

Why review is the cost center

RAND’s Where the Money Goes study of 57 large-volume productions put the split at 8% of production cost in collection, 19% in processing, and 73% in review. It was published in 2012 and is still the number cited in fee disputes, because nothing since has changed the ratio’s shape. Two automation waves have aimed at that same stage:

  • Technology Assisted Review (TAR). Predictive coding since the early 2010s. A senior attorney trains a model on a seed set, the model ranks the remaining documents by responsiveness, and low-ranked documents get lighter review or none. Court-accepted in most US jurisdictions, with fourteen years of case law behind it.
  • Generative AI review. Since 2023, LLMs handle first-pass responsiveness and privilege calls. This is now packaged product, not a pilot.

The packaging differs in a way that changes the buying decision. Relativity includes aiR for Review, aiR for Privilege, and aiR for Case Strategy in a RelativityOne subscription at no additional charge, with aiR for Case Strategy added to standard packaging on 1 July 2026. Everlaw splits the meter: single-document AI actions and the writing assistant come with the subscription, while batch actions across a document set draw on purchased credits with per-matter spend caps. Neither vendor publishes rates — both are quote-only, per-GB subscriptions.

Bundling removes per-matter AI budgeting and makes the savings invisible; metering makes AI spend legible per matter and forces a forecast. If you are the party paying the invoice, the visible meter is the better instrument. If you are the party forecasting an annual budget across dozens of matters, it is the worse one.

Specialist positioning has consolidated in the same period. Thomson Reuters retired the standalone Casetext product on 1 April 2025 and folded its technology into CoCounsel, which sells as research and drafting rather than as a review platform. Reveal has continued investing in Logikcull as the self-service tier, adding its ASK generative AI tool, Slack and Teams collection through Onna, and EU hosting in March 2026.

Do courts treat AI review differently from TAR?

No — and two Northern District of California decisions from mid-2026 are the reason to stop asking.

Generative AI review is TAR. In Schulte v. LinkedIn Corp. (N.D. Cal., No. 22-cv-00237-HSG, June 2026), Magistrate Judge Laurel Beeler let LinkedIn cull with search terms before running Relativity aiR, and refused the plaintiffs’ demand for elusion estimates and reviewer counts. Her framing was that generative AI is another predictive engine dropped into the same TAR workflow, not a new species requiring its own rules. Discovery-on-discovery stays disfavored: a party has to show a specific deficiency in the production, not a theory about the process that made it.

If you want AI transparency, negotiate it into the ESI protocol first. James v. Cerebras Systems Inc. (N.D. Cal., No. 4:25-cv-09361-AMO) shows what that looks like when a court orders it. On 7 July 2026 Magistrate Judge Robert M. Illman adopted an ESI order that requires each side to disclose the identity, version, and hosting environment of every AI system used, plus the prompts, templates, instruction sets, and parameter configurations driving it, with changes reported in redline within three business days. Validation is pinned to a 95% confidence level on null-set sampling with an elusion rate that shall not exceed 3%. The same order treats hyperlinked files as separate from document families — the requesting party may name up to 100 responsive hyperlinks, produced within 14 days at the version as sent — and requires Slack and Teams messages in searchable form with conversational context intact, batched into 24-hour periods.

Read together, the two decisions say one thing from opposite ends: courts will not police AI review on request, and they will enforce whatever you wrote into the protocol. That makes the ESI protocol negotiated in the first weeks of a matter the document where AI validation actually gets priced — a Legal Ops concern, not only an outside-counsel one.

When does a company need eDiscovery software?

For most companies: never. eDiscovery is procured per matter through outside counsel, who select the platform for that case. The company never signs a Relativity or Everlaw contract.

Companies that bring it in-house share two signals:

  1. Recurring litigation or regulatory volume. Financial services, healthcare, large tech, government contractors. The matter cadence justifies a dedicated team and a subscription.
  2. Data they do not want leaving the building. Some in-house teams run collection and processing internally, then hand a reviewed set to outside counsel.

A third signal follows from the packaging above: when AI review is bundled into a platform subscription, whoever holds the subscription captures the savings. A company with enough matter volume to fill a subscription is now choosing between capturing that itself and paying for it inside an outside-counsel rate.

  • Outside-counsel eDiscovery spend is often the largest single line item in a litigation budget. Legal spend management tooling tracks it, and alternative fee arrangements — capped per-document review — increasingly govern it.
  • Governance choices made upstream determine the downstream bill. A 90-day Slack retention policy is a multi-million-dollar decision the next time an investigation lands. That is precisely why EDRM 2.0 moves Information Governance from stage 1 to the foundation.

Common pitfalls

  • Treating the ESI protocol as boilerplate. Schulte is the price of that: once the protocol is silent on AI validation, the court will not write the term in for you afterward. Guard: negotiate AI disclosure and validation thresholds before collection begins, using the Cerebras order as the checklist — system identity and version, prompts and instruction sets, sampling confidence, elusion ceiling.
  • Buying AI review on a per-document meter without asking how a document is counted. Long attachments get split against token limits, so a 40-page exhibit can bill as several documents. Guard: get the counting rule and a per-matter ceiling in the order form, in writing, before signature.
  • Running privilege review at the same confidence bar as responsiveness. A miss on responsiveness costs a supplemental production. A miss on privilege is a waiver fight. Guard: hold privilege to its own, higher validation sample and keep a human second pass over everything the model flags as close to the line. The privilege log format you agree to shapes how much of that work is reviewable later.
  • Collecting Slack and Teams as individual messages. Stripped of thread context, a message set is unusable for review and contestable at production. Guard: collect at conversation level with the surrounding window intact — Cerebras uses 24-hour batching as the floor — and confirm your platform preserves emoji reactions and embedded media.