LLMOps : managing large language models in production / Abi Aryan.

Author/creator Aryan, Abi author.
Format Book
EditionFirst edition.
PublicationSebastopol, CA : O'Reilly Media, 2025.
Copyright Date©2025
Descriptionxiii, 266 pages : illustrations ; 24 cm
Subjects

Variant title Large language model operations
Contents Cover -- Copyright -- Table of Contents -- Preface -- Conventions Used in This Book -- O'Reilly Online Learning -- How to Contact Us -- Acknowledgments -- Chapter 1. Introduction to Large Language Models -- Some Key Terms -- Transformer Models -- Large Language Models -- LLM Architectures -- Encoder-Only LLMs -- Decoder-Only LLMs -- Encoder-Decoder LLMs -- State Space Architectures -- Small Language Models -- Choosing an LLM -- Considerations in the Selection of an LLM -- The Big Debate: Open Source Versus Proprietary LLMs -- Enterprise Use Cases for LLMs -- Knowledge Retrieval -- Translation
Contents Speech Synthesis -- Recommender Systems -- Autonomous AI Agents -- Agentic Systems -- Ten Challenges of Building with LLMs -- 1. Size and Complexity -- 2. Training Scale and Duration -- 3. Prompt Engineering -- 4. Inference Latency and Throughput -- 5. Ethical Considerations -- 6. Resource Scaling and Orchestration -- 7. Integrations and Toolkits -- 8. Broad Applicability -- 9. Privacy and Security -- 10. Costs -- Conclusion -- References -- Chapter 2. Introduction to LLMOps -- What Are Operational Frameworks? -- From MLOps to LLMOps: Why Do We Need a New Framework? -- Four Goals for LLMOps
Contents LLMOps Teams and Roles -- The LLMOps Engineer Role -- A Day in the Life -- Hiring an LLMOps Engineer Externally -- Hiring Internally: Upskilling an MLOps Engineer into an LLMOps Engineer -- LLMs and Your Organization -- The Four Goals of LLMOps -- Reliability -- Scalability -- Robustness -- Security -- The LLMOps Maturity Model -- Conclusion -- References -- Further Reading -- Chapter 3. LLM-Based Applications -- Using AI Models in Applications -- Infrastructure Applications -- Agentic Workflows -- Model Context Protocol -- Agent-to-Agent Protocol -- The Rise of vLLMs and Multimodal LLMs
Contents The LLMOps Question -- Monitoring Application Performance -- Measuring a Consumer LLM Application's Performance -- Choosing the Best Model for Your Application -- Other Application Metrics -- What Can You Control in an LLM-Based Application? -- Prompt Engineering Is "Hard" -- Did Our Prompt Engineering Produce Better Results? -- LLM-Based Infrastructure Systems Are "Harder" -- Conclusion -- References -- Chapter 4. Data Engineering for LLMs -- Data Engineering and the Rise of LLMs -- The DataOps Engineer Role -- Data Management -- Synthetic Data -- LLM Pipelines -- Training an LLM
Contents Data Composition -- Scaling Laws -- Data Repetition -- Data Quality -- A General Data-Preprocessing Pipeline for LLMs -- Step 1: Catalog Your Data -- Step 2: Check Privacy and Legal Compliance -- Step 3: Filter the Data -- Step 4: Perform Data Deduplication -- Step 5: Collect Data -- Step 6: Detect Encoding -- Step 7: Detect Languages -- Step 8: Chunking -- Step 9: Back Up Your Data -- Step 10: Perform Maintenance and Updates -- Vectorization -- Vector Databases -- Maintaining Fresh Data -- Generating the Fine-Tuning Dataset -- Automatically Generating an Instruction Fine-Tuning Dataset
Abstract Here's the thing about large language models: they don't play by the old rules. Traditional MLOps completely falls apart when you're dealing with GenAI. The model hallucinates, security assumptions crumble, monitoring breaks, and agents can't operate. Suddenly you're in uncharted territory. That's exactly why LLMOps has emerged as its own discipline. LLMOps: Managing Large Language Models in Production is your guide to actually running these systems when real users and real money are on the line. This book isn't about building cool demos. It's about keeping LLM systems running smoothly in the real world. Navigate the new roles and processes that LLM operations require Monitor LLM performance when traditional metrics don't tell the whole story Set up evaluations, governance, and security audits that actually matter for GenAI Wrangle the operational mess of agents, RAG systems, and evolving prompts Scale infrastructure without burning through your compute budget.
Bibliography noteIncludes bibliographical references and index.
ISBN9781098154202 paperback
ISBN1098154207 paperback

Availability

Library Location Call Number Status Item Actions
Joyner General Stacks QA76.9 .N38 A79 2025 ✔ Available Place Hold