Glossary
Legal AI, in plain English.
103 terms lawyers, paralegals and students meet when working with AI, from hallucinations and context windows to Opinion 512 and the EU AI Act.
ABA Formal Opinion 512
- The American Bar Association’s first formal ethics opinion on lawyers’ use of generative AI, issued July 29, 2024. It applies the Model Rules on competence, confidentiality, communication, candor, supervision and fees to AI tools. Learn more
Agent (AI agent)
- An AI system that can plan and take multi-step actions, such as searching, opening files or sending messages, towards a goal with limited human input. Agents raise heightened supervision and security questions.
Agentic workflow
- A process in which one or more AI agents carry out a sequence of tasks, often calling tools, with checkpoints for human review.
AI literacy
- The skills and understanding needed to use AI systems appropriately. Article 4 of the EU AI Act addresses AI literacy for providers and deployers.
AI system
- Under the EU AI Act, a machine-based system designed to operate with varying levels of autonomy that infers from inputs how to generate outputs such as predictions, content, recommendations or decisions. Read the full definition in the Act.
Algorithm
- A set of rules or steps a computer follows to solve a problem or complete a task.
Alignment
- Efforts to make AI systems behave in line with human intentions and values, for example by refusing harmful requests.
Annex III (EU AI Act)
- The list of high-risk use areas under the EU AI Act, including employment, education, essential services, law enforcement, migration and the administration of justice. Learn more
API
- Application Programming Interface: a way for software to connect to an AI model programmatically. API terms often differ from consumer app terms on data use.
Artificial intelligence (AI)
- Computer systems that perform tasks normally associated with human intelligence, such as understanding language, recognising patterns or making predictions.
Attention mechanism
- The technique that lets transformer models weigh which parts of the input are most relevant to each word they generate.
Automated decision-making
- Decisions made by technology without meaningful human involvement. Data protection laws such as the GDPR regulate certain automated decisions with legal or similarly significant effects.
Benchmark
- A standard test used to compare AI models. Benchmark scores rarely reflect performance on your specific legal tasks.
Bias (algorithmic bias)
- Systematic errors in AI outputs that unfairly favour or disadvantage groups, often inherited from training data.
Business associate agreement (BAA)
- A US HIPAA contract required when a vendor handles protected health information for a covered entity. Relevant when using AI on medical records in some contexts.
Candor toward the tribunal
- The duty (Model Rule 3.3) not to make false statements of law or fact to a court. Filing fabricated AI citations can breach it.
Chain-of-thought
- A prompting approach or model behaviour in which the AI works through intermediate reasoning steps before answering.
Chatbot
- A conversational interface that responds to user messages. Under the EU AI Act, people must generally be told they are interacting with an AI system unless it is obvious.
Citation hallucination
- An AI-generated reference to a case, statute or source that does not exist or does not say what is claimed. Learn more
CLM (contract lifecycle management)
- Software that manages contracts from drafting and negotiation through signature, obligations and renewal, increasingly with AI features.
Competence (technological)
- The duty (Model Rule 1.1 and Comment 8) to keep abreast of the benefits and risks of relevant technology, which ABA Opinion 512 applies to generative AI.
Confidentiality (Rule 1.6)
- The duty not to reveal information relating to a representation without informed consent or other authorisation, and to make reasonable efforts to prevent unauthorised disclosure.
Conformity assessment
- Under the EU AI Act, the process of demonstrating that a high-risk AI system meets the Act’s requirements before it is placed on the market.
Context window
- The maximum amount of text, measured in tokens, that an AI model can consider at once, including your prompt, documents and its answer. Learn more
Copilot
- A general term for an AI assistant embedded in software you already use; also Microsoft’s brand name for its AI assistants.
CRAFT framework
- A legal prompting structure taught in AI for Lawyers 2025: context and role, request, audience, format and tone. Learn more
Data processing agreement (DPA)
- A contract governing how a vendor processes personal data on your behalf, required by the GDPR and similar laws.
Data retention
- How long a vendor keeps your inputs, outputs and files. Short or zero retention reduces confidentiality risk.
Deepfake
- AI-generated or manipulated image, audio or video that appears authentic. The EU AI Act requires disclosure of deepfakes in many cases.
Deployer
- Under the EU AI Act, a person or organisation using an AI system under its authority in a professional capacity. A law firm using an AI tool is a deployer.
Diffusion model
- A type of generative model, commonly used for images, that creates content by gradually refining random noise.
Digital Omnibus on AI
- EU legislation in force from July 27, 2026 that deferred EU AI Act high-risk obligations to December 2, 2027 (Annex III) and August 2, 2028 (Annex I). Learn more
Discovery (eDiscovery)
- The identification, collection, review and production of electronically stored information in litigation, an area where AI-assisted review is long established.
Embedding
- A numerical representation of text that captures meaning, used to find similar documents in retrieval systems.
Enterprise plan
- A business tier of an AI product, usually with contractual commitments not to train on customer data, admin controls and security features.
EU AI Act
- Regulation (EU) 2024/1689, the EU’s risk-based law on artificial intelligence, in force since August 1, 2024 with obligations phased in. Learn more
Explainability
- The degree to which a human can understand why an AI system produced a particular output.
Extraction
- Using AI to pull specific data points, such as dates, parties or obligations, from documents.
Few-shot prompting
- Including a few examples of the desired input and output in a prompt so the model follows the pattern.
Fine-tuning
- Further training a pre-trained model on a narrower dataset to specialise its behaviour.
Foundation model
- A large model trained on broad data that can be adapted to many tasks; many legal AI products are built on foundation models from major AI labs.
General-purpose AI model (GPAI)
- Under the EU AI Act, a model trained on large data that can perform a wide range of tasks. GPAI model provider obligations applied from August 2, 2025.
Generative AI
- AI that creates new content such as text, images, audio or code in response to prompts.
GPT
- Generative Pre-trained Transformer, the model family behind OpenAI’s ChatGPT; also used loosely for large language models generally.
Grounding
- Connecting an AI model’s answers to specific, retrievable sources such as a legal database, so outputs can be checked.
Guardrails
- Instructions or technical controls that constrain AI behaviour, such as “do not invent citations” in a prompt.
Hallucination
- When an AI model produces information that is false or fabricated but presented confidently as fact. Learn more
High-risk AI system
- Under the EU AI Act, an AI system in an Annex I product or Annex III use area that must meet strict requirements before and after market placement.
Human in the loop
- A design in which a person reviews or approves AI outputs before they take effect.
Inference
- The process of an AI model generating an output from an input, as opposed to training.
Informed consent
- Agreement by a client after the lawyer has explained the material risks and reasonably available alternatives (Model Rule 1.0(e)). ABA Opinion 512 indicates it may be needed before inputting confidential information into some AI tools.
Input
- The prompt, documents or data you give an AI tool.
Jailbreak
- An attempt to trick an AI system into ignoring its safety rules.
Knowledge cutoff
- The date after which a model has no training data, so it may not know recent law unless connected to current sources.
Large language model (LLM)
- An AI model trained on vast text to predict and generate language. The technology behind most legal AI assistants.
Legal research platform
- A database of primary and secondary legal materials, increasingly with AI assistants grounded in that content.
Machine learning
- A branch of AI in which systems learn patterns from data rather than following explicitly programmed rules.
Metadata
- Data about data, such as authorship and edit history, which can be exposed when documents are shared or processed.
Model
- The trained AI system that produces outputs. Products often let you choose between models with different strengths.
Model Rules of Professional Conduct
- The ABA’s model ethics rules, which most US states have adopted in some form.
Multimodal
- An AI model that can handle more than one type of input or output, such as text, images and audio.
Natural language processing (NLP)
- The field of AI concerned with understanding and generating human language.
Neural network
- A computing system loosely inspired by the brain, made of layers of connected nodes that learn patterns from data.
New York Part 161
- A New York Unified Court System rule (22 NYCRR Part 161), effective June 1, 2026, permitting AI in court papers while requiring attorneys to ensure there is no fabricated material. Learn more
OCR
- Optical character recognition: converting scanned images of text into machine-readable text, often a first step before AI review.
Open-source model
- A model whose weights are publicly released so others can run or modify it. Licence terms vary.
Output
- The text or other content an AI tool produces in response to an input.
Overhead (AI costs)
- Under ABA Opinion 512, costs of general-purpose AI tools built into a firm’s software are treated as overhead rather than billed to clients.
Parameters
- The internal numerical values a model learns during training. More parameters does not automatically mean better legal performance.
Personal data
- Information relating to an identified or identifiable individual, regulated by data protection laws such as the GDPR.
Pinpoint citation
- A reference to the specific page or paragraph supporting a proposition. Always ask AI for pinpoints, then check them.
Playbook
- A set of preferred and fallback positions used to review or negotiate contracts, which some AI tools can apply automatically.
Prohibited AI practices
- Uses banned by Article 5 of the EU AI Act since February 2, 2025, such as social scoring and certain manipulative or exploitative systems.
Prompt
- The instruction or question you give an AI tool.
Prompt engineering
- The practice of designing prompts to get accurate, useful outputs. Learn more
Prompt injection
- An attack in which hidden instructions in content (such as a document or web page) manipulate an AI system into unintended actions.
Provider
- Under the EU AI Act, a person or organisation that develops an AI system or has it developed and places it on the market under its own name.
Pseudonymisation
- Replacing identifying details with placeholders so data cannot be attributed to a person without additional information. Learn more
RAG (retrieval-augmented generation)
- A technique where the system retrieves relevant documents and gives them to the model to ground its answer.
Reasonable fees (Rule 1.5)
- The duty to charge reasonable fees. ABA Opinion 512 indicates hourly billing must reflect actual time, even when AI makes work faster. Learn more
Reasoning model
- A model designed to spend more computation working through problems step by step before answering, often better at complex analysis.
Red-teaming
- Deliberately testing an AI system with adversarial inputs to find failures and risks.
Redline
- A document showing proposed changes against an earlier version.
SB 574 (California)
- California statute signed September 30, 2026, effective January 1, 2027, governing attorneys’ use of generative AI, including verification, confidentiality, court disclosure and a bar on delegating the practice of law to AI. Learn more
Self-learning tool
- An AI tool that may use inputs to improve itself, creating risk that information could surface in other users’ outputs. Opinion 512 treats these with particular caution.
Semantic search
- Search that finds results by meaning rather than exact keywords.
SOC 2
- An independent attestation framework for a service organisation’s security, availability and confidentiality controls. Type II reports cover controls over a period of time.
Standing order (AI)
- An order by an individual judge setting rules for AI use or disclosure in filings in their court.
Sub-processor
- A third party engaged by your vendor to process your data, such as the AI lab whose model powers a legal tool.
Summarisation
- Using AI to condense documents. Always ask for references back to the source.
Supervision (Rules 5.1 and 5.3)
- Duties of managerial and supervisory lawyers to ensure lawyers, staff and vendors comply with professional obligations, including in their use of AI.
Synthetic content
- Content generated or substantially altered by AI. The EU AI Act requires machine-readable marking of certain synthetic content.
System prompt
- Hidden instructions set by a product or user that shape how an AI assistant behaves in every conversation.
Temperature
- A setting controlling how random or creative a model’s output is. Lower temperature gives more predictable answers.
Token
- A small unit of text a model processes; in English, roughly three-quarters of a word on average. Learn more
Training data
- The data used to train a model. Ask vendors whether your inputs could become training data.
Transformer
- The neural network architecture behind modern large language models.
Transparency obligations (Article 50)
- EU AI Act duties, applying from August 2, 2026, to disclose AI interactions and mark or disclose certain AI-generated content.
Vector database
- A database that stores embeddings so a system can quickly find semantically similar content.
Verification
- Independently confirming that AI output is accurate, including that every authority exists and supports the point cited. Learn more
Watermarking
- Embedding signals in AI-generated content so it can be identified as synthetic.
Zero data retention
- A vendor arrangement under which inputs and outputs are not stored after processing, subject to any stated exceptions.
Zero-shot prompting
- Asking a model to perform a task without giving examples.