What Is eDiscovery? How Legal Teams Find Digital Evidence
eDiscovery is the process of identifying, collecting, reviewing and producing electronically stored information relevant to litigation or a regulatory investigation, following a structured framework known as the EDRM. If this is your organization’s first time facing an actual eDiscovery request, understanding the process before it starts genuinely changes how manageable it feels once it does.
What Is eDiscovery?
eDiscovery, short for electronic discovery, is the process of identifying, preserving, collecting, reviewing and producing electronically stored information relevant to litigation, investigation, or a regulatory request. This covers emails, documents, chat messages, database records and any other digital information potentially relevant to a specific legal matter.
Unlike cyber forensics, which focuses on deep technical reconstruction of exactly what happened on a system, eDiscovery focuses on broadly and defensibly identifying and producing relevant information a legal matter requires, a genuinely different scope even though both disciplines share overlapping technical foundations.
The Nine Stages of the EDRM, Explained
The Electronic Discovery Reference Model, EDRM, has served as the standard visual framework for eDiscovery practice since 2005, providing courts, legal teams and technology professionals a shared language for how electronic evidence moves through the process. The classic model spans nine stages: Information Governance, Identification, Preservation, Collection, Processing, Review, Analysis, Production and Presentation.
Information Governance establishes how an organization manages its data before any litigation even begins. Identification determines what information might be relevant to a specific matter. Preservation ensures that identified information is not altered or destroyed. Collection gathers that preserved information for further processing. Processing prepares collected data for review, converting and organizing it into a genuinely reviewable format. Review is where legal teams actually examine documents for relevance and privilege. Analysis identifies patterns, key facts and relationships within the reviewed material. Production delivers the relevant, non-privileged material to the requesting party. Presentation prepares that material for actual use in depositions, hearings or trial.
What’s New: Inside EDRM 2.0, the Update Currently Underway
Here is a genuinely current development worth stating precisely, since its actual status matters directly for anyone researching this topic right now. EDRM opened public comment on EDRM 2.0 on June 30, 2026, with that comment period running through July 30, 2026, marking the first substantive update to the model since it incorporated the full Information Governance Reference Model. This is not yet a finalized, adopted replacement for the classic nine-stage model covered above. It remains a draft undergoing community review, with project trustees still considering submitted feedback before determining whether further revisions are needed ahead of final publication.
Developed by roughly 150 multidisciplinary practitioners across the global EDRM community, the draft proposes four genuinely structural changes worth understanding directly. Information governance moves from a previously sidelined, often-overlooked element into the foundational layer underlying every other stage in the model. Early-phase activities, identification, preservation and collection, consolidate into a unified Data Acquisition framework positioned on the model’s left side. A standalone Disposition phase gets added specifically to address defensibly handling data after its usage in a matter concludes, an obligation the classic model never explicitly named as its own distinct stage. Analysis is repositioned as a continuous activity spanning every phase rather than a single, discrete step, reflected through bidirectional arrows showing practitioners moving iteratively between stages rather than following a strict, one-way sequence, a shift directly reflecting the growing role of analytics, machine learning and generative AI throughout the entire process. Review is specifically highlighted as the model’s nexus, the point where data volume, legal relevance and strategic decision-making genuinely converge. For any organization or legal team currently building process documentation, understanding that this update exists and remains in active development, rather than treating the classic nine-stage model as permanently fixed or already fully superseded, matters directly for staying genuinely current.
Why Review Is the Most Expensive Stage, and How AI Is Changing That
Review consistently represents the largest single cost within any eDiscovery matter, since it requires human attorneys or trained reviewers to examine every collected document individually for relevance and privilege, a genuinely labor-intensive process that scales directly with document volume rather than becoming more efficient at scale the way some other stages do.
AI, specifically through Technology-Assisted Review, is genuinely changing this cost structure by using machine learning to prioritize and predict document relevance, letting human reviewers focus their limited time on the documents most likely to actually matter rather than reviewing every single document with equal priority regardless of likely relevance. This does not eliminate human review entirely, since privilege determinations and genuinely nuanced relevance calls still require human judgment, but it meaningfully reduces the total volume requiring that human attention, directly addressing the specific cost driver that has historically made Review the most expensive stage in the entire process.
TAR vs CAL: Related, but Not the Same Thing
Technology-Assisted Review, TAR, is the broader category describing any use of machine learning to prioritize or predict document relevance during Review, a category that has existed and evolved for over a decade now. Continuous Active Learning, CAL, is a specific, more advanced protocol within that broader TAR category, distinguished by one genuinely important technical difference worth understanding directly.
Earlier TAR approaches, sometimes called TAR 1.0, trained a predictive model once on a static sample of documents, then applied that fixed model across the remaining document population without further learning. CAL instead continuously updates its own predictions in real time as reviewers code additional documents throughout the process, meaning the algorithm keeps learning and re-prioritizing what to show reviewers next based on every coding decision made so far, rather than relying on one initial training round that never adapts further. This distinction matters practically, since CAL’s continuous adaptation generally achieves stronger precision and recall over the course of a review than a static, one-time-trained model, particularly in matters where relevant documents are unevenly distributed throughout the collection. Referring to CAL and TAR as interchangeable terms technically undersells CAL’s own specific, genuine technical advancement within that broader category, precisely the kind of imprecision that can matter directly when a legal team needs to defend its chosen review methodology to opposing counsel or a court.
The Biggest Risk in Collection: Losing Metadata That Makes Evidence Inadmissible
Metadata, information about a file beyond its actual content, creation date, author, modification history, file path, often carries genuine evidentiary weight independent of the document’s own text, and improper collection methods can strip or alter this metadata in ways that undermine the evidence’s admissibility entirely.
Collecting a document by simply copying and pasting its content, forwarding an email rather than properly exporting it, or converting a file to a different format during collection, can all alter or destroy metadata that later proves essential to establishing exactly when a document was created, who actually authored it, or whether it was modified after a relevant date. This is precisely why forensically sound collection tools and documented methodology matter directly, capturing files in their native format with metadata genuinely intact, rather than collection methods that happen to preserve the visible content while quietly destroying the metadata a court or opposing counsel may later specifically need to verify the evidence’s authenticity and timeline. A collection process that looks complete based on document content alone can still represent a genuine evidentiary failure if the underlying metadata never survived that collection process intact.
Information Governance: The Stage That Starts Before Litigation Ever Happens
Information Governance sits at the very start of the EDRM specifically because it addresses how an organization manages its data long before any specific litigation or investigation ever begins, encompassing retention policies, data classification and defensible deletion practices that directly shape how manageable a future eDiscovery request eventually becomes.
An organization with genuinely mature information governance, clear retention schedules, organized data classification, defensible, documented deletion of data no longer needed, faces a considerably smaller, more manageable universe of information when an actual eDiscovery request eventually arrives, compared to an organization that has never actively managed its own data lifecycle at all. This is precisely why Information Governance’s positioning at the model’s starting point, and its elevated, foundational role within the proposed EDRM 2.0 update covered above, reflects genuine, practical reality rather than an abstract, theoretical placement: the work done here, years before any litigation is ever anticipated, directly determines how expensive and complicated every subsequent stage eventually becomes.
What Does This Look Like for a Business Facing Its First eDiscovery Request?
Issue a legal hold notice immediately to every relevant custodian, formally instructing them to preserve potentially relevant information and specifically suspend any routine deletion or auto-purge processes that would otherwise destroy it. Identify the specific individuals, systems and data sources genuinely likely to hold information relevant to the matter, rather than attempting to preserve everything indiscriminately across your entire organization.
Engage a professional eDiscovery or forensic collection service specifically for the actual collection stage, rather than attempting informal, DIY collection methods risking exactly the metadata loss covered above, since a first-time eDiscovery matter is precisely the wrong moment to discover collection mistakes have compromised your evidence. Document every step taken throughout the process, since a defensible, documented methodology matters as much for a business facing its first request as it does for an established litigation support team handling matters routinely. Cyber Security Solutions Ltd helps businesses navigate exactly this first-time scenario, since the specific mistakes made during an unfamiliar, first eDiscovery request are almost always preventable with proper guidance from the very first legal hold notice onward.
Conclusion
eDiscovery follows a genuinely structured process, and understanding where your organization currently sits within that process, and that the framework itself is actively evolving through EDRM 2.0, matters directly whether you’re preparing proactively or facing your first request right now. Start by confirming your own information governance practices before litigation ever makes that question urgent. To get help navigating an eDiscovery request or building stronger information governance beforehand, visit cybersecuritysolutionsltd.com for expert support from Cyber Security Solutions Ltd.
FAQs
eDiscovery is the process of identifying, preserving, collecting, reviewing and producing electronically stored information relevant to litigation, investigation or a regulatory request, following the structured EDRM framework that has served as the standard for this process since 2005.
Information Governance, Identification, Preservation, Collection, Processing, Review, Analysis, Production and Presentation. Each stage prepares information for the next, from establishing how data is managed before litigation through to presenting evidence at trial.
EDRM 2.0 opened for public comment on June 30, 2026, and remains a draft under active community review, not yet a finalized replacement for the classic model. Proposed changes include information governance as a foundation and a new Disposition phase.
TAR, Technology-Assisted Review, is the broader category of using machine learning to prioritize document relevance. CAL, Continuous Active Learning, is a specific, more advanced TAR protocol that continuously updates predictions in real time as reviewers code documents.
Improper collection methods, like copying content or converting file formats, can strip or alter metadata showing creation date, author and modification history. This metadata often carries genuine evidentiary weight, and its loss can undermine evidence admissibility entirely.
Issue a legal hold notice immediately to relevant custodians, identify specific individuals and systems likely to hold relevant information, and engage a professional collection service rather than attempting informal collection that risks compromising metadata and evidence integrity.
