What is a large language model (LLM)?
An LLM is a software system trained on vast collections of text—books, web content, code and research papers—to understand and generate human language. Built on transformer architecture, an LLM processes language by simultaneously weighing relationships between words and phrases across the entire input, rather than reading left to right in sequence. This enables handling complex, context-dependent queries at scale.
Unlike rule-based natural language processing (NLP) systems, which match inputs against predefined patterns and fail when language falls outside those patterns, LLMs generate responses based on statistical relationships learned during training. This distinction matters more than it might appear: rule-based NLP fails predictably—you can audit the failure, trace it to a missing rule and fix it—while LLMs fail probabilistically, meaning they produce fluid and plausible but sometimes factually inaccurate outputs.
The transformer architecture enables LLMs to handle tasks that would require separate, hand-coded systems in traditional NLP: e.g., summarizing a 200-page contract, answering questions across a knowledge base or generating code from a plain-language description. While such breadth is genuinely useful, it brings governance challenges and operational risks.
How LLMs work: Tokens, training and inference
Tokenization
Before a model reads any text, that text is broken into tokens—units that may correspond to whole words, word fragments or punctuation marks, depending on the model's vocabulary (e.g., "contract renewal" might become three tokens). Each LLM has a maximum number of tokens it can process per interaction.
Pre-training
A model processes enormous text datasets and learns to predict what token is likely to follow a given sequence, adjusting its internal parameters across billions of iterations, resulting in a model that has absorbed broad patterns of language, reasoning and factual association without being explicitly programmed with any of it.
Inference
It predicts the next token, then the next, building the output sequentially based on learned probabilities and the provided context. Inference latency—the time between submitting a prompt and receiving a complete response—is a procurement-grade constraint for real-time applications (e.g., live customer-facing chatbot).
LLM types: Foundation, fine-tuned and small language (SL) models
While it's common to view foundation models as preferred—with fine-tuned and SL models seen as less expensive options—that's inaccurate and leads to the deployment of expensive, ungovernable and underperforming systems. Let's compare:
| Dimension | Foundation Model | Fine-tuned Model | Small Language Model (SLM) |
|---|---|---|---|
| Definition | Broad-corpus LLM, no task-specific tuning | Foundation model adapted to domain-specific data | Compact LLM optimized for efficiency |
| Training scope | General—diverse text at scale | Domain-specific—layered on pre-trained base | Narrow—task or domain constrained |
| Parameter scale | Larger | Larger, with domain-adapted weights | Fewer parameters |
| Enterprise fit / when to use | Prototyping, broad NLP tasks, general-purpose assistants | Regulated industries, high-stakes domain tasks, compliance-sensitive workflows | Latency-sensitive applications, cost-constrained deployments, on-prem requirements |
| Trade-off | High cost; governance overhead; weaker domain precision | Requires quality domain data; additional training investment | Narrower capability ceiling; task scope must be well-defined before deployment |
Example (reference points, not recommendations) | GPT-4 | BloombergGPT (fine-tuned model trained on financial text) | Mistral 7B |
LLM risks
Three unsolvable risks define enterprise LLM governance; each requires a design decision, not mitigation.
Hallucinations LLMs generate outputs based on learned probabilities rather than verified facts, producing confident, fluent and sometimes factually incorrect responses. The business consequence isn't occasional inaccuracy—it's that the errors may be indistinguishable from correct output without external verification. All models hallucinate, so the design questions are:
- Can my workflow detect hallucinations?
- Are the consequences of an undetected error acceptable?
Data privacy
When enterprise data is submitted to a third-party LLM via API, it's used in model training and exposed to other users. While strong governance requires auditable, stable system behavior, LLMs are updated continuously to avoid degradation and model irrelevance. There's no formula for balancing governance rigor and model relevance, so this risk must be managed rather than eliminated.
Bias
LLMs unavoidably absorb and amplify the distributional biases present in the internet-scale text on which they're trained, including underrepresentation of certain languages, industries and demographic groups.
Enterprise LLM use cases
- Healthcare document summarization
Providers can generate structured summaries from unstructured clinical notes, reducing clinicians' documentation time while preserving contextual details that keyword extraction misses. - Manufacturing code generation and knowledge management
Firms can automate the generation of boilerplate integration code between operational technology systems and enterprise software, where the logic is repetitive but the syntax varies by platform. They can also surface relevant technical documentation and maintenance procedures from repositories where terminology varies across product generations and regions. - Finserv customer support automation and contract analysis
LLM-powered support systems can handle variations in account inquiries that would exhaust rule-based chatbot logic, while routing edge cases to human agents based on confidence thresholds. During due diligence, legal teams can flag non-standard indemnification clauses across large contract portfolios.
How to evaluate and select an LLM for your enterprise
- Accuracy benchmarks: No standardized method for generating domain-specific benchmarks exists. You need to build task-specific evaluation sets from your own data, so accuracy is partly a post-procurement discovery.
- Latency: Define your latency ceiling before testing, not after and measure response time under realistic concurrent loads, not in isolated single-query tests.
- Cost per token: Token-based pricing compounds quickly at enterprise scale. Model costs against expected query volume before procurement. A lower per-token cost on a less capable model may be acceptable for well-scoped workflows.
- Compliance: Identify which data residency, logging and model update policies apply to your regulatory environment before evaluating any vendor. A model that doesn't meet them simply isn't a candidate.
- Vendor ecosystem fit: Assess integration with your existing infrastructure, including authentication systems, data pipelines and monitoring tooling. A model that requires significant custom integration work carries a hidden cost.
- Healthcare document summarization
In evaluations, CIOs focus on cost per token and compliance, IT architects on latency under load and ecosystem integration.









