LLMOps : managing large language models in production / Abi Aryan.
| Author/creator | Aryan, Abi author. |
| Format | Book |
| Edition | First edition. |
| Publication | Sebastopol, CA : O'Reilly Media, 2025. |
| Copyright Date | ©2025 |
| Description | xiii, 266 pages : illustrations ; 24 cm |
| Subjects |
| Variant title | Large language model operations |
| Contents | Cover -- Copyright -- Table of Contents -- Preface -- Conventions Used in This Book -- O'Reilly Online Learning -- How to Contact Us -- Acknowledgments -- Chapter 1. Introduction to Large Language Models -- Some Key Terms -- Transformer Models -- Large Language Models -- LLM Architectures -- Encoder-Only LLMs -- Decoder-Only LLMs -- Encoder-Decoder LLMs -- State Space Architectures -- Small Language Models -- Choosing an LLM -- Considerations in the Selection of an LLM -- The Big Debate: Open Source Versus Proprietary LLMs -- Enterprise Use Cases for LLMs -- Knowledge Retrieval -- Translation |
| Contents | Speech Synthesis -- Recommender Systems -- Autonomous AI Agents -- Agentic Systems -- Ten Challenges of Building with LLMs -- 1. Size and Complexity -- 2. Training Scale and Duration -- 3. Prompt Engineering -- 4. Inference Latency and Throughput -- 5. Ethical Considerations -- 6. Resource Scaling and Orchestration -- 7. Integrations and Toolkits -- 8. Broad Applicability -- 9. Privacy and Security -- 10. Costs -- Conclusion -- References -- Chapter 2. Introduction to LLMOps -- What Are Operational Frameworks? -- From MLOps to LLMOps: Why Do We Need a New Framework? -- Four Goals for LLMOps |
| Contents | LLMOps Teams and Roles -- The LLMOps Engineer Role -- A Day in the Life -- Hiring an LLMOps Engineer Externally -- Hiring Internally: Upskilling an MLOps Engineer into an LLMOps Engineer -- LLMs and Your Organization -- The Four Goals of LLMOps -- Reliability -- Scalability -- Robustness -- Security -- The LLMOps Maturity Model -- Conclusion -- References -- Further Reading -- Chapter 3. LLM-Based Applications -- Using AI Models in Applications -- Infrastructure Applications -- Agentic Workflows -- Model Context Protocol -- Agent-to-Agent Protocol -- The Rise of vLLMs and Multimodal LLMs |
| Contents | The LLMOps Question -- Monitoring Application Performance -- Measuring a Consumer LLM Application's Performance -- Choosing the Best Model for Your Application -- Other Application Metrics -- What Can You Control in an LLM-Based Application? -- Prompt Engineering Is "Hard" -- Did Our Prompt Engineering Produce Better Results? -- LLM-Based Infrastructure Systems Are "Harder" -- Conclusion -- References -- Chapter 4. Data Engineering for LLMs -- Data Engineering and the Rise of LLMs -- The DataOps Engineer Role -- Data Management -- Synthetic Data -- LLM Pipelines -- Training an LLM |
| Contents | Data Composition -- Scaling Laws -- Data Repetition -- Data Quality -- A General Data-Preprocessing Pipeline for LLMs -- Step 1: Catalog Your Data -- Step 2: Check Privacy and Legal Compliance -- Step 3: Filter the Data -- Step 4: Perform Data Deduplication -- Step 5: Collect Data -- Step 6: Detect Encoding -- Step 7: Detect Languages -- Step 8: Chunking -- Step 9: Back Up Your Data -- Step 10: Perform Maintenance and Updates -- Vectorization -- Vector Databases -- Maintaining Fresh Data -- Generating the Fine-Tuning Dataset -- Automatically Generating an Instruction Fine-Tuning Dataset |
| Abstract | Here's the thing about large language models: they don't play by the old rules. Traditional MLOps completely falls apart when you're dealing with GenAI. The model hallucinates, security assumptions crumble, monitoring breaks, and agents can't operate. Suddenly you're in uncharted territory. That's exactly why LLMOps has emerged as its own discipline. LLMOps: Managing Large Language Models in Production is your guide to actually running these systems when real users and real money are on the line. This book isn't about building cool demos. It's about keeping LLM systems running smoothly in the real world. Navigate the new roles and processes that LLM operations require Monitor LLM performance when traditional metrics don't tell the whole story Set up evaluations, governance, and security audits that actually matter for GenAI Wrangle the operational mess of agents, RAG systems, and evolving prompts Scale infrastructure without burning through your compute budget. |
| Bibliography note | Includes bibliographical references and index. |
| ISBN | 9781098154202 paperback |
| ISBN | 1098154207 paperback |
Availability
| Library | Location | Call Number | Status | Item Actions |
|---|---|---|---|---|
| Joyner | General Stacks | QA76.9 .N38 A79 2025 | ✔ Available | Place Hold |