Get in Touch

Course Outline

Day 1: Build the Foundation — Ingest, Search, Retrieve

Module 1: The Legal Engineer’s Landscape

  • Learning objectives — understand the role, where AI fits into legal work, and the two critical risks that permeate the field.
  • Topics
    • The legal-engineer role and why it is currently in high demand
    • Where AI fits: eDiscovery, review, contracts, research, investigations; the EDRM model explained simply
    • Build vs. buy decisions
    • The two persistent risks: confidentiality/privilege and defensibility

Module 2: Legal Data Is Messy — Ingestion and Extraction

  • Learning objectives — handle the reality of legal data at scale.
  • Topics
    • 1,400+ file types, email and PST formats, scanned paper, load files (.dat/.opt); relevant embedded metadata
    • Text extraction (Tika), OCR, and deduplication strategies
  • Lab: FreeEed Ingestion — build an ingestion pipeline over a deliberately messy document set (email/PST, scans, load files)

Module 3: Search and Retrieval — The Foundation

  • Learning objectives — build the core eDiscovery primitive: finding anything inside everything.
  • Topics — full-text search and indexing (Solr/Lucene); relevance, metadata, and date filtering; searching across OCR’d content
  • Lab: eDiscovery Search — index a corpus and execute real eDiscovery-style searches, including within OCR’d scans

Module 4: RAG for Legal Documents — with Citations

  • Learning objectives — build RAG over legal documents that cites its sources.
  • Topics
    • Why retrieval, not fine-tuning, is preferred for sensitive material — the model never consumes the documents directly
    • Chunking, embeddings, and above all, citations/provenance
    • Multi-document and thread summarization
  • Lab: Legal RAG with Citations — build a RAG Q&A over a document set that answers with source citations

Day 2: Make It Private, Defensible, and Shippable

Module 5: Privacy, Privilege, and Local Serving — The Privilege Trap

  • Learning objectives — keep legal data local and certify its security.
  • Topics
    • Where data actually goes when interacting with cloud AI
    • Privilege waiver, duty of competence, and the “private” spectrum (contractual vs. physical)
    • Morgan v. V2X case study and why local deployment is court-defensible
    • Serving local models (Ollama/vLLM) and monitoring outbound traffic
  • Lab: Local Model + Egress Proof — run a local model end-to-end and prove, via monitoring, that no data exited the system

Module 6: Defensible AI Review

  • Learning objectives — measure and document an AI review so it withstands scrutiny.
  • Topics
    • Court-admissible metrics: recall, elusion, precision, ground-truth validation; TAR/active learning
    • Transparency (why did it code this document?) and reproducibility — pinning the model, fixing settings, logging everything
    • The “defensible case snapshot” allowing someone to re-run your review a year later with identical results
  • Lab: Defensible Review — measure an AI review against a blind ground truth and produce a reproducibility bundle

Module 7: Ship It — Workflow, Private Deployment, and Governance

  • Learning objectives — assemble pieces into a workflow, deploy it privately, and score its performance.
  • Topics
    • A multi-step legal workflow (ingest → search → summarize → review → produce) with human-in-the-loop
    • Private/on-premises deployment essentials (containerization; keeping data in-house)
    • AI governance for legal, briefly, and scoring the system with SAIS-100 (the Elephant Scale Secure AI Score)
  • Lab: Score and Package — wire a multi-step workflow, score it with SAIS-100, and package it for private deployment

Capstone (integrated across Day 2)

  • Build a private, defensible legal-AI application end to end — ingest a messy corpus, search it, answer questions over it with citations using a local model, measure a defensible review, and package it for private deployment.
  • Participants leave with a portfolio project that represents the work of a legal engineer.

Optional Day 3 / Advanced Modules (delivered as a 3rd day or modular series)

  • Investigations: Entities, Relationships, and Timelines — extract people/orgs/dates, reconstruct email threads, build chronologies, map near-duplicates and document lineage. Lab: build a timeline and entity/relationship view.
  • Agentic and Multi-Step Legal Workflows (deep) — richer orchestration, contract analysis, multi-doc synthesis, tool use, and guardrails as a design principle. Lab: build a multi-step workflow with a human checkpoint.
  • Deployment at Scale — on-premises and appliance deployment, distributed processing for large volumes, regulated environments (CJIS, government, higher-ed), hardware sizing. Lab: containerize and scale a processing job across workers.
  • Governance and Compliance Deep-Dive — the AI-regulation landscape (100+ US state AI laws, the EU AI Act), audit requirements, and a full SAIS-100 governance audit. Lab: audit a legal-AI system against a governance/defensibility checklist.

Requirements

  • Proficiency in Python and basic APIs
  • Helpful: User-level familiarity with LLMs (no machine learning background required; mental models will be developed)
  • No legal background required — necessary legal concepts are taught within context

Audience

  • Software and AI engineers transitioning into the legal tech sector
  • Engineers at legal-tech companies needing deeper domain-specific knowledge
  • Technically inclined legal, eDiscovery, or information governance professionals who wish to build rather than just purchase tools
  • Anyone aiming for a “legal engineer” or “AI legal engineer” role
 14 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories