The XCP-ng Developer's Handbook

Table of Contents

Introduction

What this book is about

Running multiple services in production raises a fundamental tension. The simplest approach is to dedicate a physical machine to each service, which guarantees isolation: services cannot affect each other. But in practice, each machine would use only a small fraction of its available resources, and the cost of buying and maintaining a large number of machines quickly becomes prohibitive. The natural response is to share physical resources across services. But sharing in turn raises isolation and security concerns: how do you prevent services from affecting each other when they run on the same machine? Getting the most out of your hardware while keeping your services well isolated from each other is the core problem that virtualization addresses.

It does so by running services inside virtual machines instead of directly on physical hardware. From the point of view of the software running inside it, a virtual machine looks and behaves like a real computer, with its own operating system, memory, storage, and network interfaces. The physical machine hosting the virtual machines is called the host. The virtual machines themselves are called guests. Note that these are relative terms: a virtual machine can itself run other virtual machines, in which case it is simultaneously a guest from the point of view of the physical machine, and a host from the point of view of the virtual machines running inside it.

The idea of running services inside virtual machines turns out to have profound consequences for the way infrastructure is built and operated. It also turns out to be surprisingly difficult to implement well. XCP-ng is one of the platforms that implements it, and this book is about how it does so.

XCP-ng has its roots in XenServer, a virtualization platform developed by Citrix. In 2018, following changes in Citrix's licensing policy, XCP-ng was created as a fully open-source fork of XenServer. The two projects continue to share a common codebase, and contributions flow between them. This history explains why some existing documentation refers to XenServer: much of it remains relevant to XCP-ng, and we will point to it where appropriate.

Who this book is for

This book is written for experienced software developers whose goal is to become contributors to XCP-ng, but who have no background in virtualization. Prior experience with virtual machines or containers as a user may help, but is not required. No prior experience with production infrastructure or systems administration is assumed either.

Parts I, II and III of this book, which cover virtualization concepts and the architecture of XCP-ng, will be useful to anyone seeking to understand the platform, regardless of whether they intend to contribute. Part IV, which covers the contribution process, is specifically aimed at developers who want to get involved in the project.

What this book does not cover

This book is not a replacement for existing documentation. What it provides is a map: a coherent, progressive path through the project, from the first principles of virtualization to the practical details of contributing to XCP-ng, building the foundations that other documentation tends to assume. Throughout the book, we refer to existing documentation where relevant. Readers who want to go deeper into specific topics may find that existing documentation is easier to approach after reading this book.

Beyond this, the following topics are outside the scope of this book:

  • Advanced network and storage configuration.
  • Performance tuning and optimization.
  • XenOrchestra, except where necessary to understand the boundary between what XCP-ng and XenOrchestra each handle.
  • Third-party integrations and plugins.
  • Security hardening.
  • System design and distributed systems architecture. Virtualization can be seen as one of the building blocks that distributed systems are built upon, but the broader questions of how to design systems that scale, tolerate failures, and remain performant under load are outside the scope of this book. Chapter 1 touches on orchestration as one of the challenges of running services at scale, and Chapter 4 situates XCP-ng in the broader ecosystem of infrastructure tools, both of which will help clarify where this book's scope ends and distributed systems architecture begins.

These topics are either covered by existing documentation or would require books of their own.

How this book is organized

The book is divided into four parts and a conclusion.

Part I provides a tour of the virtualization landscape. Chapter 1 presents the challenges inherent in deploying and maintaining services at scale. Chapter 2 then presents the different answers available to address these challenges: containers, emulators, and virtualization of different kinds, and positions XCP-ng among them.

Part II looks at XCP-ng from the outside. Chapter 3 traces its history, from the Xen research project to the 2018 fork from XenServer, and explains how this history continues to shape the project today. Chapter 4 situates XCP-ng in the broader ecosystem of virtualization management tools, container orchestration platforms, and cloud providers. Chapter 5 then walks through the installation process, explores what the system looks like once installed, and shows how to perform basic operations on virtual machines. This gives the reader a concrete foundation before we open the hood.

Part III looks at XCP-ng from the inside. We describe its overall architecture, introduce its main components, and devote a chapter to each of them.

Part IV explains how to effectively contribute to XCP-ng. We explain how the project is organized, how to set up a development environment, and how to test your changes.

The conclusion reviews what has been covered and provides guidance on how to go further, with pointers to resources for staying up to date with the project and the broader virtualization ecosystem.

Part I: The Virtualization Landscape

Before diving into XCP-ng itself, we need to understand the landscape it belongs to. This part is organized in two chapters. Chapter 1 identifies the challenges that arise when running services at scale: resource efficiency, isolation, portability, resilience, and others. Chapter 2 then surveys the main approaches that have been developed to address these challenges, from virtual environments and containers to virtualization, and positions XCP-ng among them.

By presenting all the challenges first, we would like to encourage the reader to keep all the constraints at play in mind while examining each solution, making it easier to understand not just what each approach does, but also why it exists and what trade-offs it makes.

Chapter 1: The challenges of running services at scale

Running services at scale means operating in an environment where the number of services, users, or operational constraints has grown beyond what can be managed naively. This chapter lays out the constraints that arise in such environments.

As discussed in the Introduction, running different services on different machines is inefficient, but running them on the same one raises isolation and security concerns. Other constraints include keeping them available despite hardware failures and software updates, managing multiple environments from development to production, and orchestrating the whole system, meaning coordinating tasks like deploying, monitoring, and restarting large numbers of services and machines.

These constraints do not exist independently of each other: a decision made with resource efficiency in mind will have consequences for resilience; a choice driven by security requirements will affect how environments are managed. They thus need to be considered together, which is why we present them all before examining any solution. The set of criteria that emerges from this chapter will help situate each of the approaches presented in Chapter 2.

Resource efficiency

The simplest way to allocate resources (CPU, RAM, storage, network bandwidth and more generally any device) to a service is to dedicate a physical machine to it. Although this has the appeal of simplicity (the service has exclusive access to all the resources of the machine), it comes with an intrinsic inefficiency that we will now explain.

A machine must be capable of handling peak load. As brief as the peak may be, the machine must be sized to handle it, which means that for the rest of the time, a fraction of its CPU or memory remains unused. This is not a failure of planning: it is a structural property of dedicating a machine to a single service.

Beyond this waste of resources, a dedicated machine carries operational costs that are easy to underestimate. On the hardware side, components fail and must be repaired or replaced, and when the service grows and the machine can no longer keep up, capacity must be added, for instance by adding RAM, replacing disks with larger ones, or adding network interfaces. On the software side, the operating system and all installed software must be kept up to date. And when the machine reaches the end of its life, it must be decommissioned: any remaining data must be migrated, the machine must be wiped, and then either recycled or donated.

But an organization running tens or hundreds of services on dedicated machines would face these costs for each machine in the fleet, multiplying both the inefficiency and the operational burden.

And so does the environmental cost. Indeed, each step of each machine's life cycle does have an environmental impact. Manufacturing consumes raw materials, including rare earth elements whose extraction carries its own increasing environmental cost. Running consumes electricity, and so do the cooling systems that keep datacenters from overheating, which in many cases also consume significant quantities of water. Finally, recycling itself consumes energy and cannot even be complete: some components must be discarded, representing an irreversible loss of the physical resources that went into making them. At the scale of tens or hundreds of machines, these costs add up to an environmental impact that is no longer negligible.

Isolation

As established in the previous section, running one service per machine is not viable. The only practical approach is thus to run multiple services on the same machine, which implies sharing its resources. This in turn translates into two constraints.

The first is the equitable sharing of resources. No service should be able to lessen the others in terms of access to the machine's resources.

The second is the isolation of data and privileges. A service should not be able to read the data of another service running on the same machine, nor benefit from its privileges.

Availability

From a user's perspective, a service is either available or it is not. When it is not, whatever the reason, the user cannot do what they came to do. Availability is therefore not a nice-to-have: for any service that people depend on, it is a hard requirement.

Yet a service running on a single machine cannot fulfill this requirement: any failure, whether a hardware fault, a software crash, or a necessary maintenance operation, interrupts it, which requires redundancy: running multiple instances of the service on multiple machines, so that if one machine fails, the others take over without interruption.

Redundancy, however, introduces a new constraint. Running multiple instances of the same service on different machines only works if those instances are consistent with one another. A service deployed differently on each machine, with different configurations or different versions of its dependencies, will behave unpredictably. The ability to deploy a service reliably and consistently across multiple machines is therefore a prerequisite for redundancy.

This need extends beyond production. Development and testing environments must be close enough to the production environment for tests to be meaningful, yet distinct from it to avoid any risk of disrupting the service. This again requires the ability to deploy the service consistently in multiple contexts.

Finally, keeping a service available over time raises the question of updates. Any service in production will inevitably need to be updated, whether to add new features, fix bugs, or address security vulnerabilities. But updating a service, whether the service itself or its dependencies, risks introducing regressions or, if done carelessly, causing the very interruption that redundancy was designed to prevent. How to update a service running on multiple machines, progressively and without interruption, is a problem that takes us directly to the question of orchestration.

Orchestration

Most of what has been discussed so far has concerned a single service. In practice, however, an organization runs tens or hundreds of services, each with its own life cycle: initial deployment, successive updates, monitoring, restarts when it fails, and eventual decommissioning. Managing these life cycles manually is already demanding for a handful of services, and even then it is error-prone. At the scale of tens or hundreds of services, each running on multiple machines, it is simply not an option.

The situation is further complicated by the fact that services often depend on one another. A service that relies on another cannot be deployed or updated independently of it. Coordinating deployments and updates across interdependent services is exponentially more complex than managing each service in isolation.

The same applies to development. Developing and testing a service requires an environment that is close enough to production for the results to be meaningful, which adds yet another context in which the full set of dependencies must be deployed coherently.

Orchestration is the name given to this problem: how to coordinate the development, testing, deployment, monitoring, and life cycle of a large number of services running on a large number of machines, reliably and without constant manual intervention.

Summary: the set of constraints

As has been seen, running one service per machine is not an option. Beyond its economic and environmental costs, it is simply not viable in practice. The only remaining option is thus to run multiple services on the same machine, which raises the need for both equitable access to its resources and isolation between services (no service should be able to access the data or privileges of another).

Moreover, keeping services available requires redundancy, which in turn requires the ability to deploy services consistently across multiple machines and in multiple environments. And as the number of (often interdependent) services grows, so does the need to coordinate their deployment, monitoring, and life cycles without constant manual intervention.

Throughout this chapter, we have identified a set of constraints that arise when running services at scale. As announced in the introduction, it is against this set of constraints that we will situate each of the approaches presented in the next chapter. To help the reader keep them in mind while reading, here they are in synthesized form:

  1. Run multiple services on the same machine, with equitable access to its resources: CPU, memory, storage, network bandwidth, and devices.
  2. Prevent any service from accessing the data or privileges of another.
  3. Deploy services across multiple machines so that they remain available despite hardware failures and can be updated consistently and without interruption.
  4. Keep development, testing, and production environments consistent with one another while keeping them distinct.
  5. Coordinate the deployment, monitoring, and life cycles of many often interdependent services without constant manual intervention.

Chapter 2: Addressing the challenges

The previous chapter identified a set of constraints that arise when running services at scale. This chapter examines the main approaches that have been developed to address them. For each approach, we first describe what it is and how it works, then situate it against the constraints identified in the previous chapter.

Not all approaches address all constraints. This is not a shortcoming: satisfying each constraint has a cost, in terms of complexity, resources, or trade-offs with other constraints, and different approaches make different choices about which costs are worth paying. Understanding these choices is as important as understanding the approaches themselves.

The approaches examined here range from language-level virtual environments and OS-level containers to full emulation and virtualization of different kinds. Among them, the mechanisms that operating systems provide for scheduling, isolation, and access control occupy a particular place: they represent the earliest systematic response to constraints 1 and 2, and they also introduce concepts that are necessary for understanding the approaches that follow. XCP-ng is a virtualization platform, and by the end of this chapter the reader will understand where it sits among these approaches and why.

Language-level isolation: virtual environments

We have not yet asked how a service is installed on a machine. The natural starting point is the system package manager: apt on Debian-based systems, dnf on Red Hat-based systems, and their equivalents elsewhere. A package manager not only installs a package but also, before doing so, identifies and installs a list of packages it depends on. For each of these dependencies, there may be an additional constraint on the version: a package may require another at version 2.0 or above, or between versions 1.4 and 2.0, or exactly at version 3.1. The package manager's job is to find a set of package versions that satisfies all these constraints simultaneously and to install them.

Such a system of constraints typically has many solutions. The package manager picks one, and the result is a coherent set of packages whose components are expected to work together. From a user's perspective, this is a desirable property: the system stays current as packages are updated and constraints evolve, the package manager may find a new consistent set, and coherence is maintained at each step.

For a developer working across multiple projects, though, this model shows its limits. Two projects may require incompatible versions of the same library: no single version can satisfy the constraints of both. Since the system package manager allows only one version of a given library to be present at a time, no solution exists that satisfies both projects simultaneously. There is also a second pressure: language ecosystems tend to move faster than distributions, with new package versions published directly by their authors rather than going through the slower process of distribution packaging. Both factors contributed to the emergence of language-specific tools that address the problem within their own scope.

These tools emerged pragmatically, within specific communities, to solve immediate needs. Their language-specificity is both a strength and a limit: a tool designed for a given ecosystem can integrate intimately with its conventions, package metadata, and build tooling, in a way that a language-agnostic tool could not. But it can only ever solve the problem within that ecosystem.

The approach these tools share is straightforward: a dedicated directory that holds its own copy of the packages installed for a given project, combined with a set of environment variable modifications that redirect the language's tools toward that directory instead of the global one. A program launched with the appropriate environment will find its dependencies in that directory rather than in the system-wide store. The packages installed in one environment are invisible to another, and conflicts between projects become impossible in principle, within the bounds of what the language's tooling manages.

Like system package managers, these tools express dependencies as constraints on version numbers and must find a solution that satisfies them all. Unlike system package managers, many of them record that solution in a lock file: a precise, complete record of the exact version of every dependency as resolved at a given point in time. A lock file transforms a set of constraints, which may have many solutions, into a single fixed solution that can be reproduced exactly across machines and over time. Tools like poetry, stack, and uv generate such lock files as a matter of course.

These tools differ in what they expose to the user. An environment can be local to a single project, or shared across several; some tools offer one mode, some the other, and some both. npm in JavaScript and stack in Haskell treat the environment as a property of the project, as does Cargo in Rust. nvm in JavaScript and rbenv in Ruby offer named environments that can span multiple projects. venv and virtualenv in Python leave the choice to the user. opam switch in OCaml also supports both modes: switches can be global and shared across projects, or local to a specific project directory.

The conda Python tool occupies a distinct position in this landscape. Unlike the tools described above, it can manage dependencies written in other languages, including C libraries that Python packages may depend on. This makes it more general than the other tools described here, but it remains an ecosystem-specific solution. As we will see in the next section, Nix and Guix generalize this further in a principled way, handling dependencies of any kind, for any project, within a single unified framework.

Tools such as Snap, Flatpak, and AppImage take a different approach: rather than isolating dependencies in a separate directory, they bundle them directly inside the package. Unlike the ecosystem-specific tools described above, they are not tied to any particular language. This eliminates conflicts with system libraries for the bundled application, but these tools are designed primarily for deploying desktop applications, not services, and they fall outside the scope of what we examine in this chapter.

In terms of the constraints identified in the previous chapter, these ecosystem-specific tools address a narrow slice. The first constraint asks that multiple services be able to run on the same machine with equitable access to its resources: these tools do not control resource usage in any way. The second constraint asks that no service be able to access the data or privileges of another. These tools offer little here: two services running in different environments share the same operating system and the same file system outside their respective directories. More fundamentally, the isolation they provide rests entirely on a cooperative assumption: every tool in the ecosystem must agree to respect the environment variables that redirect package lookups. A package that hardcodes a path to a system library, or a build script that bypasses the package manager entirely, will silently break out of the isolation with no external mechanism to prevent it. The boundary exists only because every participant agrees to observe it. As we will see, this assumption persists in different forms through the approaches that follow, and addressing it fully requires mechanisms of a different kind entirely.

The third constraint concerns deployment across multiple machines and availability despite hardware failures. These tools are silent on availability and fault tolerance, though the use of lock files does facilitate consistent deployment across machines, which is a partial contribution to this constraint. The fifth constraint, concerning the coordination of many interdependent services, is likewise not addressed.

The fourth constraint is where these tools have something genuine to offer: the requirement that development, testing, and production environments remain consistent with one another. By isolating dependencies on a per-project basis and, where lock files are used, fixing an exact solution to the dependency constraints, these tools make it possible to reproduce a known package configuration across machines and over time. This reproducibility, however, extends only to the packages the tool manages: dependencies at the system level remain outside its reach. That boundary is precisely what Nix and Guix set out to eliminate.

Declarative packaging and configuration: Nix and Guix

As has been seen, system package managers have been designed to maintain a coherent set of packages across installations and updates: they have not been designed to address reproducibility concerns. Neither do they permit the coexistence of two versions of the same package. The tools examined in the previous section address these issues only partially: they are mostly confined to a single ecosystem, which means they are not general enough to cover all dependencies in a uniform way. And even for the dependencies they do cover, nothing prevents a package from relying on anything other than its declared dependencies.

The Nix package manager was designed to overcome these limitations. Its central idea is that each package is identified by a hash derived from all of its inputs: its source code, build instructions, and build-time and run-time dependencies. Readers familiar with Git will recognize the structure: just as a commit hash uniquely identifies a point in a repository's history by capturing both the state of the tree and its ancestry, a Nix package hash uniquely identifies a build result by capturing everything that went into producing it. Nix installs each package in its own directory, as ecosystem-specific tools do; what distinguishes it is that the directory name is derived from the package's hash, so any difference in inputs, however slight, results in a distinct package in a distinct location.

This more complete and rigorous management of dependencies does not only allow packages to be distinguished more precisely from one another: because each package is fully determined by its inputs, the build process becomes much more deterministic, producing the same output from the same inputs on any machine. This property is what we call reproducibility. Since this is a primary concern for Nix, it ensures that the package's build process can only access what has been explicitly declared as an input, so that only its declared inputs can influence the result. This isolation is enforced by mechanisms provided by the operating system, which we will examine later in this chapter. It is worth pausing on this point. In the previous section, isolation relied on a cooperative assumption: every package had to agree to respect the conventions that made isolation possible. Here, for the first time, we encounter a constraint that is enforced mechanically rather than observed by convention. A package's build process cannot reach outside its declared inputs, not because it is expected to refrain from doing so, but because it is made impossible. This isolation applies only during the package's build stage: once a package has been built, the resulting binaries run without any restriction imposed by Nix.

This approach, too, has its limits: not all dependencies can be expressed as Nix inputs. The kernel is the most significant example: it is provided by the host system and lies outside what Nix can describe or control. A build that depends on kernel-specific values, like the output of the uname command, will therefore produce different results on machines with different kernel versions, even if all its other inputs are identical. NixOS addresses this by extending Nix's declarative approach to the configuration of an entire system, including its kernel: it thus becomes possible to express all aspects of a machine's configuration as Nix inputs, and to reproduce it exactly on another machine. Guix and Guix System are the GNU project's equivalents of Nix and NixOS respectively. Where Nix uses a dedicated language designed specifically for package description, Guix uses Guile, a general-purpose Scheme implementation: package descriptions are Scheme values that can be computed, which means the full range of Scheme's factorisation and reuse mechanisms are available to facilitate the development and maintenance of package descriptions and build recipes.

Against the constraints identified in the previous chapter, Nix and Guix have little to say about the first: they exercise no control over how services share CPU time, memory, or any other resource.

The second constraint asks that no service be able to access the data or privileges of another. The cooperative assumption is not eliminated here but reduced in scope. At the package build stage, isolation is enforced: a package cannot silently depend on anything other than its declared inputs. Once the package has been built, though, the resulting binaries run without any restriction imposed by Nix: nothing prevents them from reading files outside their environment, opening network connections, or interacting with other running programs. The activation of a Nix environment, like that of an ecosystem-specific tool, relies on environment variable adjustments that any program is free to ignore. The boundary that Nix draws is precise and reliable at the package build stage, but once the resulting binaries are running, the cooperative assumption remains.

The third constraint concerns deployment across multiple machines and continued availability despite hardware failures. NixOS and Guix System address the first part: because a system configuration is fully declared, it can be reproduced exactly on another machine, which makes consistent deployment easier. The second part, availability despite hardware failures, however remains outside their scope.

The fourth constraint is where Nix and Guix offer their most substantial contribution. Because every dependency is a declared input with a fixed identity, the reproducibility they provide extends to the entire user-space environment. NixOS and Guix System extend this further by bringing the kernel configuration within the scope of what is declared, reducing though not eliminating the remaining sources of variance: even with a fully declared system configuration, factors such as differences in underlying hardware or non-deterministic behaviour in the kernel itself remain outside the reach of any declarative tool.

The fifth constraint, concerning the coordination of many interdependent services, is not addressed by these tools either.

OS foundations: scheduling, memory isolation, and access control

Operating systems have long provided mechanisms to address the first two constraints identified in the previous chapter, those that arise when sharing a machine between multiple services. The first is ensuring equitable access to its resources. The second is the need to prevent any service from accessing the data or privileges of another.

In response to the first constraint, the main mechanism provided by operating systems is scheduling, which governs equitable access to the CPU. Equitable access to other resources such as memory, storage, and network bandwidth is not addressed at the level of operating systems. The following section will introduce a more complete mechanism: containers.

In response to the second constraint, operating systems provide two distinct mechanisms. The first is memory isolation, which prevents a service from reading or writing the memory of another. The second is access control through users, groups, and permissions, which governs access to files and devices. Access to network resources is not addressed at this level either. Here again, it is containers that will fill in the gaps.

This section does not follow the evaluative structure used everywhere else in this chapter: rather than presenting a single approach and then situating it against all five constraints, it examines a set of related mechanisms that together constitute the earliest systematic response to constraints 1 and 2, and that also provide the conceptual foundations for the approaches examined in the following sections. It is also the first time in this book where we move beyond pure software: some of the mechanisms we examine here (scheduling and memory isolation) rely on features provided by the processor itself.

  • Scheduling

    For a long time, sharing a machine meant sharing a single processor. Even though this is no longer true today, it is in this context that the mechanisms we examine here were developed.

    A processor can execute only one service at a time. So if several services are to make progress simultaneously, the only possible way is to run them in turn. For that to work, it must be possible to interrupt a service and resume it without the service being affected by the interruption. This in turn requires that when a service is interrupted, its state is saved so that it can be restored when the service resumes, allowing it to continue exactly where it left off.

    However, the fact that it must be possible to interrupt a service does not mean that this ability should be given to anyone. Indeed, if any service could interrupt any other, equitable access to the processor could not be guaranteed. Hence the need for an arbitrator, which is the sole entity entitled to decide which service has access to the processor at any given time. Such an arbitrator is called a scheduler. It maintains a list of running services, each entry of which is called a process. For processes that are not currently executing, this entry holds their state as saved at their last interruption, allowing them to be resumed transparently. Each entry also holds whatever information the scheduler needs to decide which process should execute next. For instance, it holds the frequency and duration of processor access allocated to that process, parameters that the scheduler can adjust over time based on the recent behavior of the process. Most schedulers also allow a priority to be assigned to each process, which also influences the frequency and duration of processor access it receives.

    Having an arbitrator is a necessary condition for guaranteeing equitable access to the processor, but it is not sufficient. The scheduler must also be effectively capable of preempting other processes, that is, of taking the processor away from them without them being able to prevent it. Such a capability can only be conferred by the processor itself. This requires hardware support: the processor must provide a privileged mode, reserved for the scheduler, and a restricted mode in which services execute, with the hardware itself enforcing the boundary between them.

    At this stage, two important remarks are due. On the one hand, the fact that processors provide several execution modes is not an implementation detail. On the contrary, it is a cornerstone without which there could be no scheduling and thus no real possibility to run several services on the same machine. On the other hand, our presentation based on two execution modes, a protected one for the scheduler and a restricted one for the services, is deliberately simpler than the reality, for pedagogical reasons. As will be seen later in this chapter, virtualization complicates this simple picture. And virtualization is only one reason why the reality is more complex.

    The following subsections will provide further examples of components that require privileged execution. The collection of all components that run in privileged mode is generally called a kernel. That is why the privileged mode in which the kernel executes is more commonly called kernel mode. Similarly, the restricted mode in which processes execute is more commonly called user mode.

    With this new terminology, preemption can be described as a transition from user mode to kernel mode initiated by the hardware. Not all such transitions, however, are initiated by the hardware. They can also be initiated by the process through a system call: a request to the kernel to perform a privileged operation.1

  • Memory Isolation

    Early operating systems gave each process direct access to physical memory. Any process could read or write any address, including addresses belonging to other processes. This made isolation between services impossible to guarantee.

    Virtual memory was introduced to solve this problem while preserving the addressing model that programs rely on. Each process is given its own address space, which it perceives as a contiguous and exclusive range of addresses. The addresses a process uses are virtual addresses, which do not correspond directly to physical addresses.

    Memory is divided into fixed-size blocks called pages. The kernel maintains, for each process, a page table mapping its virtual addresses to physical addresses. This table contains entries only for pages that have been allocated to that process. A process can only form virtual addresses that fall within its own address space: it has no knowledge of the physical addresses used by other processes, and no way to reach them. Isolation is thus guaranteed by construction, not by explicit checks at each memory access.

    The translation from virtual to physical addresses is performed at every memory access by a unit integrated into the processor called the MMU (Memory Management Unit). The kernel configures the page tables; the MMU uses them to perform translations. This is another instance of the cooperation between software and hardware seen in the previous subsection: the kernel cannot enforce memory isolation by software alone, and the MMU cannot operate without the page tables the kernel maintains.

    Virtual memory will reappear in a more complex form when we examine virtualization.

  • Access Control

    Virtual memory prevents a process from accessing the memory of another, but it says nothing about files and devices. A process that cannot access another process's memory can still open its files, overwrite its configuration, or access a device it should not. Addressing constraint 2 in full thus requires a separate mechanism, one that governs access to files and devices.

    For such a mechanism to be effective, it must be impossible for any process to bypass it. This is achieved by requiring all access to files and devices to go through the kernel, via system calls. The kernel is thus the natural enforcer of access control for files and devices.

    The problem is then one of identification and authorization. To decide whether a given process may access a given resource, the kernel must first be able to identify both who owns the process and who owns the resource.

    Unix addresses this through a model known as discretionary access control (DAC), built on three concepts: users, groups, and permissions. Each process runs on behalf of a user, and each file belongs to a user. This makes it possible to express ownership, but ownership alone is not enough: it is also useful to be able to control access at a coarser granularity than individual users. This is the role of groups: a file belongs to a group as well as to a user, and all members of that group are subject to the same access level for that file.

    Access itself, however, is not a single indivisible property. It can be split into three atomic ones: reading a file's contents, writing to it, and executing it. Permissions define, for each of these properties, whether it is authorized. They do so for three scopes separately: the owner of the file, the members of its group, and all other users.

    Finally, one aspect of isolation that the DAC model does not address is the structure of the filesystem itself. A process with appropriate permissions can navigate the entire filesystem. Nothing in the permission model restricts which part of it a process may explore.

    The chroot system call, introduced in Unix in 1979, was an early response to this concern. It changes the root directory of a process to a specified directory, confining the process to a subtree of the filesystem. Files outside that subtree become invisible to the process: it cannot name them, and therefore cannot access them regardless of permissions.

    chroot reflects a preoccupation with filesystem isolation that predates most of the approaches examined in this chapter. It is also a conceptual ancestor of the namespace mechanisms that will be examined in the following section.

OS-level Isolation: Containers

We now introduce containers and evaluate them against the five constraints identified in the previous chapter. Containers rely on three mechanisms (namespaces, cgroups, and images) that we describe in turn, before explaining how containers are built on top of them.

Namespaces. The previous section introduced chroot, which confines a process to a subtree of the filesystem. A process confined in such a way cannot name files outside that subtree, and therefore cannot access them. However, chroot only isolates the filesystem. Other aspects of a process' execution environment, such as process identifiers, network interfaces, and users, remain shared with the rest of the system. A process confined by chroot can still see other processes and send them signals, access any network interface of the host, and share its view of users and groups. Namespaces generalize the principle of chroot to many aspects of a process' execution environment.

Here are a few examples of namespaces available on Linux. A mount namespace gives a process its own table of mounted filesystems, independently of the host's. (In particular, a process can mount any filesystem as its root, which means that chroot can be seen as a special case of what a mount namespace provides.) A PID namespace gives a process its own set of process identifiers: processes outside the namespace are invisible to it, and even the identifiers it sees bear no relation to those used on the host. A network namespace gives a process its own set of network interfaces and routing tables. In each case, the isolation is structural rather than based on explicit access checks: the entities a process cannot access are simply absent from its view of the system. This is the same principle we observed with virtual memory, where a process cannot reach the memory of another simply because that memory does not appear in its address space.

Cgroups. Namespaces control what a process can see. However, they provide no mechanism to express or enforce limits on what it can consume. Except CPU time which is covered by scheduling, nothing prevents a process running in its own set of namespaces from exhausting most of the resources available to the host, starving other processes of them. This is what cgroups, short for control groups, address. They make it possible to limit and account for the resource usage of a process or a group of processes: CPU time (complementing the scheduling discussed earlier with precise quotas and usage accounting), memory, network bandwidth, and others. Where namespaces provide isolation, cgroups provide containment. Like scheduling and virtual memory, both namespaces and cgroups are implemented as kernel mechanisms; unlike them, though, they require no dedicated hardware support.

Images. Namespaces and cgroups together define the boundaries of a process' execution environment. It still remains to be specified what will run in that execution environment. Just as a program is an executable that can be run by an operating system, an image is an executable that can be run inside an execution environment. Such an image consists of both a read-only filesystem that embeds the service to be run and its dependencies, and an entry point: the program that will actually be executed when the image is run. Because an image is self-contained, it can be run on any host, independently of what is installed on it.

Containers. Launching an image involves creating a mount namespace and mounting its filesystem as the root of that namespace, creating additional namespaces to isolate various aspects of its execution environment, such as its view of processes, network interfaces, and users, and creating cgroups to bound its resource consumption. Its entry point is then executed: it runs as a process in that confined environment. Just as running a program yields a process, running an image yields a container. This is not merely an analogy: at the system level, running an image amounts to running its entry point, a program, which effectively yields a process. Just as a program can be run several times, each run yielding a distinct process, likewise an image can be run several times, each run yielding a distinct container. And, as processes can be started, stopped, resumed, and terminated, so can containers: they, too, have a life cycle.

The software that made containers popular is Docker, and it remains the most widely used containerisation platform. It makes two design choices in the name of simplicity: it centers on a daemon, and it runs that daemon as root. The first choice creates a single point of failure: if the daemon stops, all containers become unreachable. The second has a cost in terms of security: any user authorized to communicate with the daemon thus has effective root access to the host, and can, for instance, ask it to mount the host's root filesystem into a container started from an otherwise ordinary image, hence gaining full access to the host's whole filesystem.

An alternative that addresses these drawbacks is Podman. At the cost of a more complex implementation, it departs from both of Docker's choices: it requires no daemon, and containers can be run without any root privileges, a mode known as rootless containers.

Now, turning to the constraints of Chapter 1: containers address constraint 1 through cgroups. Where the scheduling discussed in the previous section already protected the CPU, cgroups extend this protection to memory, network bandwidth, and other resources, preventing any service from monopolizing them.

Constraint 2 is addressed through namespaces: since the resources of a container are simply not visible from others, it becomes impossible for any service to access either the data or the privileges of that container.

Constraint 3 is only partially addressed: images ensure that a service can be deployed consistently across multiple hosts, which is a necessary condition for distributed deployment, but keeping services available despite hardware failures, and updating them without user-visible interruption, requires coordinating containers across a cluster of machines, which is beyond what containers themselves provide; it is the role of cluster management tools such as Kubernetes or Docker Swarm.

Constraint 4 is addressed by images: because an image is immutable and self-contained, images provide the same degree of environment consistency as the declarative approaches examined in Section 2 of this chapter. Here too, the main remaining source of variation is the host's kernel.

Finally, constraint 5 is not addressed at all by containers themselves. Coordinating the deployment, monitoring, and life cycles of many interdependent services requires tooling that operates above the level of individual containers. Tools such as Docker Compose handle this at small scale, while tools such as Helm address it at the scale of a cluster of machines.

Emulators

All the approaches examined so far were developed with the constraints of the previous chapter as a central concern. Emulators were not: they were designed to serve different needs, and whatever they have to say about those constraints is incidental. We examine them here because they occupy an important place in the landscape of systems software, and because some of the concepts they rely on, in particular the distinction between host and guest, are the same ones we will encounter again when we turn to virtualization.

The first need is compatibility, understood broadly. Software written for a system that no longer exists must somehow be made to run on current hardware. But compatibility has a forward-looking face as well: a developer may want to write software that is independent of any particular platform, meaning any combination of hardware and operating system, so that it can run anywhere without modification. These two motivations are different, but they call for the same mechanism.

The second need is control. Running software in an environment that the developer fully controls, where execution can be observed, interrupted, and replayed, is valuable both for debugging and for testing. This need overlaps more directly with some of the constraints of the previous chapter, though it was not originally framed in those terms.

An emulator addresses these needs by reproducing the behavior of a target platform on a host platform. The guest program runs as though it were executing on the target platform; the host platform runs the emulator, which decodes and executes each guest instruction on the guest program's behalf. Familiar examples include FCEUX, which emulates the Nintendo Entertainment System and runs on Linux, Windows, and macOS; RetroArch, a frontend that bundles emulators for a wide range of consoles and runs on the same platforms; and DOSBox, which emulates a complete x86 PC including its processor, memory, and peripherals, and runs on Linux, Windows, and macOS. In each case the emulated system is the guest and the machine running the emulator is the host. These emulators serve the compatibility need: they make it possible to run software written for hardware that no longer exists or is no longer practical to use.

QEMU, in its emulation mode, serves both needs. It can emulate a wide range of processor architectures, making it possible to run software compiled for one architecture on a machine with a different architecture. It can also emulate a given architecture on itself, which is useful for testing and debugging: a program running under QEMU has no direct access to the host hardware, and its execution can be observed and interrupted at will.

The Java Virtual Machine follows a related but distinct logic. Rather than reproducing the behavior of an existing physical system, the JVM defines an abstract machine: a set of instructions, called bytecode, that correspond to no real processor, for which Java programs are compiled. Any conforming JVM implementation will execute them identically, regardless of the underlying host platform. This is what "virtual" means in "Java Virtual Machine": the machine exists only as a specification, not as physical hardware. The JVM was introduced in 1995 with the slogan "write once, run anywhere," and it serves the portability face of the compatibility need: write for the abstract machine, and the implementation takes care of the rest. This sense of "virtual machine" is different from the one we will encounter in the next section, where a virtual machine is a reproduction of a real physical hardware environment rather than an abstract specification.

The idea behind the JVM is older than it might seem. In 1979, Infocom designed the Z-machine, a specification for an abstract processor for which all of their interactive fiction games were compiled. A Z-machine interpreter was then written for each platform Infocom wanted to support, making the games themselves fully portable. Any conforming Z-machine interpreter will produce the same results, because the guarantee rests not on a particular implementation but on the rigor of the specification. The JVM applied the same principle to general-purpose programming sixteen years later.

Not all systems that model another system's behavior aim for this kind of faithful reproduction. A simulator models some aspects of a system's behavior for the purpose of observation or analysis, without attempting to reproduce the system completely or execute its actual code. The iOS simulator included in Xcode is a clear example. An iPhone has sensors, an accelerometer, a gyroscope, a compass, that have no physical counterpart on a Mac. The simulator does not reproduce this hardware; instead, it provides a software interface through which the developer can inject artificial sensor data, for instance simulating a rotation or a change of orientation using controls in the Xcode interface. The application under development interacts with this interface as though it were talking to real hardware. The goal here is not faithful reproduction but a controlled environment in which the application can be developed and tested without requiring a physical device.

Finally, Wine occupies a different position in this landscape. It allows programs compiled for Windows to run on Linux, but it does not emulate the Windows hardware environment or interpret Windows executable code. Instead, it provides reimplementations of the Windows API libraries, such as kernel32.dll and user32.dll. When a Windows program calls a function from one of these libraries, Wine's implementation of that function translates the call into the equivalent Linux system call. The Windows program's own code runs natively on the host processor throughout. This is why Wine's name is a recursive acronym for "Wine Is Not an Emulator": compatibility is achieved not by reproducing a hardware environment but by reimplementing the software interface that the program depends on. The result is that Windows programs run at native speed under Wine, but only if Wine implements the API calls they make.

As noted at the outset, the tools examined in this section were not designed as responses to the constraints of the previous chapter, and it would be misleading to evaluate them as if they were. Nevertheless, some observations are worth making. On constraint 1, running multiple services on the same machine with equitable access to resources, emulators bring nothing specific: multiple emulators can run concurrently on the same host, but this is simply because they are ordinary processes managed by the host operating system's scheduler. On constraint 2, preventing any service from accessing the data or privileges of another, emulators offer a form of isolation as a byproduct of how they work: a guest program can only access the resources that the emulator exposes to it. The quality of this isolation depends entirely on the emulator's implementation and on whatever isolation mechanisms the host operating system makes available; it is not guaranteed by construction. On constraint 3, deploying services across multiple machines so that they remain available despite hardware failures, emulators have nothing to offer regarding availability or seamless updates; they do, however, contribute to portability, as discussed above. On constraint 4, keeping development, testing, and production environments consistent, emulators are genuinely useful: because an emulator reproduces a target platform deterministically, a program that runs correctly under the emulator in development can be expected to behave the same way in production, regardless of differences in the host hardware. On constraint 5, coordinating the deployment, monitoring, and life cycles of many interdependent services, emulators have nothing to offer: this constraint is simply outside their scope.

Machine Virtualization

This section occupies a particular place in this chapter. The approaches examined so far were chosen because they address, each in its own way, the constraints identified in the previous chapter. Machine virtualization is examined here for the same reason, but also because it is the central subject of this book: XCP-ng is a virtualization platform, and the concepts introduced in this section will be used throughout the rest of this book.

We start by defining what machine virtualization is, then turn to introducing hypervisors, before evaluating machine virtualization against the constraints identified in the previous chapter.

  • What Virtualization Is

    The previous chapter identified, among other constraints, the need to run multiple services on the same machine with equitable access to its resources and without any service being able to access the data or privileges of another. The approaches examined so far constitute successive responses to these constraints, each offering stronger isolation than the previous one. Virtual environments and Nix and Guix rest on a cooperative assumption: the isolation they provide depends on every component agreeing to observe it. Containers lift this assumption partially, enforcing isolation through kernel mechanisms rather than convention. The isolation they provide, however, remains bounded by a fundamental constraint: all containers running on a given machine share that machine's kernel. A vulnerability in the kernel exposes all of them simultaneously, and a flaw in the namespace mechanisms can allow a container to escape to the host.

    Machine virtualization lifts this assumption entirely, but at a cost. Rather than sharing a kernel, each service runs inside its own complete operating system, with its own kernel. The isolation between services is therefore as strong as the isolation between two separate machines, which is why each such isolated operating system instance is called a virtual machine. The cost of this strength is significant: where a conventional system runs one kernel per machine, machine virtualization requires at least one kernel per service2.

    Once this principle is established, the question arises of how to execute the instructions of a guest system. One answer is to interpret them one by one: this is precisely what the emulators examined in the previous section do, with the advantage in terms of flexibility regarding processor architectures and the performance cost we had identified. Machine virtualization chooses a different answer: to have the host hardware execute those instructions directly. This requires that the host and the guest share the same processor architecture, but it eliminates the interpretation overhead entirely.

  • Hypervisors
    • Motivation And Definition

      A guest operating system running on a host machine expects to run in kernel mode, just as it would on real hardware. On the host, however, kernel mode is already occupied by the host kernel, and the hardware itself enforces this: no other component can enter kernel mode without the host kernel's involvement. Two operating systems therefore cannot both run in kernel mode on the same machine at the same time.

      When we examined how multiple processes share a machine, we found that this situation calls for a privileged component that stands above the processes and that the hardware itself protects from being usurped: the kernel. The same logic applies here: running multiple kernels on the same machine requires a privileged component that stands above them in the same way. This component is called a hypervisor.

    • Dealing With Privileged Instructions

      Although Popek and Goldberg established in 1974 that a more privileged execution mode would be necessary for hypervisors to function properly, it is only some years later that hardware actually provided such a mode. Two software solutions were developed in the interim, each working within the constraints of the existing two-mode execution model.

      The first software solution, binary translation, works as follows: the hypervisor inspects the instructions of the guest kernel before they are executed and replaces privileged instructions on the fly with equivalent sequences that can be executed without conflict. From the guest's perspective, everything appears to work as expected. This approach requires no modification to the guest system and thus makes it possible to virtualize any operating system, including one that has no knowledge of being virtualized. The cost of this generality is performance: the hypervisor must inspect and potentially rewrite instructions before every execution. VMware popularized this technique in its early products.

      The second software solution, paravirtualization, takes a different approach. Rather than correcting privileged instructions after the fact, the guest operating system itself is modified to replace them with explicit calls to the hypervisor, known as hypercalls. Because the guest cooperates actively with the hypervisor, the overhead of binary translation is eliminated, and performance is significantly closer to that of a native system. The cost of this efficiency is flexibility: only guest systems available in a version designed to run in a virtualized environment can be paravirtualized. Xen, introduced in 2003, made paravirtualization its primary approach.

      When hardware support eventually arrived, with Intel's VT-x and AMD's AMD-V extensions, it introduced the third execution mode that Popek and Goldberg had identified as necessary. With this mode available, the hypervisor could run at a privilege level above kernel mode, eliminating the conflict between host and guest kernels entirely. This made it possible to run unmodified guest operating systems efficiently: the performance cost of binary translation and the flexibility constraint of paravirtualization were both overcome.

    • Design Choices

      Historically, the design of hypervisors has been shaped by two concerns: their position relative to the host hardware, and the functionality they must provide, which depends on what is being virtualized.

      On the first concern: the description given earlier of the hypervisor as the most privileged component on the machine naturally leads to a first kind of implementation: a hypervisor that runs directly on the hardware, with no host operating system beneath it. This is called a type 1 hypervisor. In practice, however, a different need also arises: adding virtualization capabilities as a layer on top of an existing system, without replacing its foundations, in the same way that a new software component can be installed on a running system without disrupting it. A hypervisor that meets this need by running on top of an existing host operating system, as an application among others, is called a type 2 hypervisor. A type 2 hypervisor does not itself occupy the highest privilege level: it is the host kernel that remains the most privileged component and that provides the underlying mechanisms the hypervisor depends on. Because there is no standard interface between type 2 hypervisors and kernels, each type 2 hypervisor must arrange for the kernels it wants to run on to expose the primitives it needs, typically by installing kernel modules. It is these kernel modules that use the hardware virtualization extensions and execute in the most privileged mode; the rest of the hypervisor runs in user space, handling tasks such as device emulation and management.

      On the second concern: the functionality a hypervisor must provide depends on what it is asked to virtualize. Two broad use cases have emerged. The first is desktop virtualization: virtualizing complete interactive working environments, with graphical interfaces, audio devices, USB peripherals, and other hardware that users interact with directly, and where the responsiveness of the interactive experience is a primary concern. The second is server virtualization: running services on dedicated hardware, for instance in a data center, where the primary concerns are to abstract the underlying hardware so that services remain independent of the physical machines they run on, and to maximize isolation, density (that is, the number of virtual machines that can run simultaneously on a given host), and availability. These two use cases pull in opposite directions: desktop virtualization must expose the host hardware as faithfully and richly as possible, while server virtualization seeks to abstract and hide it.

      The traditional type 1 / type 2 distinction tended to reinforce this separation: type 2 hypervisors such as VMware Workstation, VirtualBox, and Citrix Virtual Apps and Desktops were adopted primarily for desktop use, while type 1 hypervisors such as VMware ESXi, Xen, and Microsoft Hyper-V were adopted primarily for server use.

      KVM illustrates how a different approach can partly transcend these distinctions. Rather than being a standalone hypervisor of either type, KVM is a kernel module that turns the Linux kernel itself into a hypervisor. It runs directly on the hardware, giving it the control and isolation of a type 1 hypervisor, while remaining integrated with a full general-purpose operating system, which gives it access to the rich ecosystem of Linux drivers and tools. This combination allows it to serve both desktop and server virtualization depending on the context, and it has become the foundation of many modern virtualization stacks, including those used by major cloud providers.

    • The Xen Hypervisor

      Among the type 1 hypervisors mentioned above, Xen deserves closer examination, as it is the foundation on which XCP-ng is built.

      In Xen's terminology, every virtual machine running on the hypervisor is called a domain. Given the privilege level at which the hypervisor executes, Xen takes the approach of minimizing the code that runs at that level, moving out of the hypervisor everything that can be moved out. Indeed, every line of code that runs at the highest privilege level is a potential source of vulnerabilities, and keeping that code minimal thus reduces the attack surface.3 The functionality that does not need to reside in the hypervisor is gathered into a dedicated privileged virtual machine, called Dom0 (domain zero), which is the first domain to start when the system boots. The hypervisor grants Dom0 direct access to the hardware and ensures that no other domain has this permission. Dom0 also has the ability to create, configure, and destroy other domains. The remaining domains, which have none of Dom0's privileges, are called DomU (unprivileged domains). This approach, which keeps the hypervisor minimal by delegating responsibilities to Dom0, is sometimes described as a thin hypervisor.

      This design is a step towards a more secure architecture: by moving privileged code out of the hypervisor, it reduces the attack surface at the highest privilege level. Dom0, however, still concentrates a great deal of trust, and if it is compromised, an attacker gains control over the entire system. This has motivated work on breaking Dom0 itself into smaller components, each holding only the privileges strictly necessary for its function, an approach sometimes called Dom0 disaggregation.

      Finally, the distinction between full virtualization and paravirtualization, which shaped the early history of machine virtualization, has become less sharp over time. When hardware support arrived with VT-x and AMD-V, it became possible to run unmodified guest operating systems efficiently, without relying on binary translation. Xen, which had introduced paravirtualization as its primary approach, adopted hardware support as well, and modern versions of Xen use both approaches depending on the circumstances: paravirtualization where the guest system supports it, and hardware-assisted full virtualization otherwise. QEMU follows a similar logic in a different register: in its emulation mode, it interprets each guest instruction, at the cost of performance; when used together with KVM, it delegates the execution of guest instructions to the hardware and handles only device emulation itself, combining the flexibility of emulation with the performance of hardware-assisted virtualization.

  • How Virtualization Addresses The Constraints Of Chapter 1

    Machine virtualization addresses the first constraint, equitable access to resources, at a different level than the approaches examined previously. Operating system scheduling and container cgroups both operate at the level of individual processes: they control how processes share CPU time, memory, and other resources. Machine virtualization operates at the level of entire virtual machines: the hypervisor allocates resources to each virtual machine as a whole. The mechanisms differ, but the result is comparable: no single service can monopolize the resources of the host.

    The four other constraints all benefit from the same central property of machine virtualization: each virtual machine runs its own kernel, making it a self-contained and portable unit whose isolation from other virtual machines is as strong as the isolation between two separate physical machines. This property comes at a cost, however: this approach requires running one kernel per virtual machine, which represents a significant overhead in memory and storage compared to the approaches examined previously.

    For the second constraint, preventing any service from accessing the data or privileges of another, this property means that there is no shared kernel through which one service could reach another. The isolation is not absolute, however: the hypervisor itself is a shared component, and depending on the chosen virtualization solution, Dom0 or the host kernel may also be shared. A vulnerability in any of these components could potentially be exploited to cross the boundary between virtual machines.

    For the third, fourth, and fifth constraints, the central observation is that because a virtual machine encapsulates a complete environment, kernel included, it can be treated as an opaque and portable unit by any tool that operates above it. Unlike the approaches examined previously, which remain dependent on the host kernel and must delegate the management of that dependency to coordination tools, virtual machines make the problem disappear entirely: any host running the same hypervisor can run the virtual machine.

    For the third constraint, deploying services across multiple machines and keeping them available despite hardware failures, a virtual machine can therefore be moved from one host to another without any dependency on the host's kernel. Keeping services available despite hardware failures and updating them without interruption, however, still requires coordination tools that operate above the level of individual virtual machines.

    For the fourth constraint, keeping development, testing, and production environments consistent with one another, machine virtualization goes further than any of the approaches examined previously: it eliminates the main remaining source of variance that was identified for containers and for Nix and Guix, namely the dependency on the host kernel.

    For the fifth constraint, coordinating the deployment, monitoring, and life cycles of many interdependent services, machine virtualization does not provide a solution by itself, just as the approaches examined previously do not. Coordination tools that operate above the level of individual virtual machines are required. Because virtual machines are fully encapsulated units, such tools can treat them as opaque and portable objects, providing a better foundation for building them than the approaches examined previously.

Part II: XCP-ng from the outside

Chapter 3: A brief history of XCP-ng

The Xen research project

XenServer: Xen goes commercial

The 2018 fork: birth of XCP-ng

An ongoing relationship with Citrix

Chapter 4: XCP-ng in the ecosystem

Management and orchestration: XenOrchestra, OpenStack, CloudStack

Container orchestration: Kubernetes

Configuration management: Ansible, Puppet

Cloud platforms: AWS, Azure, Google Cloud

Where XCP-ng fits in

Chapter 5: XCP-ng: a Linux distribution

The installer

What is on the ISO image

What is on the installed system

Chapter 6: Setting up a working environment

Installing XCP-ng on one or more hosts

Setting up a pool

Basic virtual machine operations

Part III: XCP-ng from the inside

Chapter 7: Overall architecture

Xen: the hypervisor

Dom0 and DomU

XenStore

XAPI

How these components fit together

Chapter 8 and beyond: one chapter per component

Part IV: Contributing to XCP-ng

Chapter x: Organization of the project

Repositories

Contribution process

Chapter x+1: Setting up a development environment

Chapter x+2: Testing your changes

Conclusion

What has been covered

How to go further

Footnotes:

1

The set of operations that require a system call has two origins. Some operations are reserved to kernel mode by the processor itself – the corresponding instructions simply cannot be executed in user mode. For others, the choice of making them a kernel prerogative is a design decision of the operating system – it is not imposed by the processor. In both cases, the kernel remains the mandatory intermediary: a process running in user mode can only access the underlying resources to the extent that the kernel permits.

2

As will be explained soon, the exact number depends on the type of hypervisor.

3

This principle is shared by microkernel operating systems such as GNU Hurd, which keep the kernel itself minimal and delegate most services to user-space components. The motivation is the same: a smaller trusted computing base is easier to reason about and to secure.

Author: Seb Hinderer

Created: 2026-07-09 Thu 14:10

Validate