Once again, we find ourselves with a new type of CPU that is unlike anything seen on the store shelves. After all, this is another 7th-generation console that reflects an obsessive need for innovation, a peculiar trait during that era.
Introduction
Before we discuss the architecture, I'll start with a bit of history to bring you up-to-speed. The following paragraphs focus on the business aspect of the Xbox 360's CPU, whose sequence of events you may find amusing, to say the least.
I'll try to keep it short so we can focus on the main topics of this series, but if in the end, you are interested in more, you may enjoy a book called 'The Race For A New Game Machine' which is written by the former executives at IBM.
From complier to ruler

The beginning of Microsoft's presentation at E3 2005, where Robbie Bach, J Allard and Peter Moore unveiled the Xbox 360 .
After enjoying the surprising success of the original Xbox, it was time for Microsoft to work on the successor. The company started by looking for vendors that could face off against Sony's upcoming technology, however, unlike their previous development, Microsoft now enjoyed the upper hand in forthcoming negotiations.
To put things in context, back when the original Xbox project was still at an early stage, neither Intel nor Nvidia were willing to share their intellectual property with Microsoft. This decision limited Microsoft's capacity to mould Nvidia's or Intel's chips for the specific needs of the Xbox.
For instance, the security subsystem that protected the console against the execution of unauthorised code was implemented outside these two critical chips. This made it vulnerable to snooping attacks that eventually paved way for the execution of Homebrew and piracy. Moreover, Microsoft did not control the manufacturing stage either, so the production of Xbox systems was at the mercy of Intel's and Nvidia's supply.
Well, now that Microsoft had gained more leverage in the console market, they weren't willing to give away those rights anymore.
Sharing common problems

The original Xbox's CPU (2001). Designed and manufactured by Intel.

Concept of an homogeneous CPU. This is what Microsoft had in mind for their next CPU.
As with any other company in the computer business, the innovation crisis of the early noughties affected both Microsoft and Sony indiscriminately. The difference, however, was that the two placed their bets on different designs for their CPUs. The original Xbox relied on popular off-the-shelf stock (Intel's Pentium III) with slight customisations, this was a single-core CPU extended with vectorised instructions and a sophisticated cache design. On the other side, Sony's vectorised venture (the Emotion Engine) consisted of a low-end CPU surrounded by proprietary but potent assistants.
For their new generation, Microsoft settled for a conservative design combined with some experimental ideas. According to their checklist, the new CPU would be multi-core (that is, a chip housing many CPU cores in a symmetrical and homogeneous layout) and the supplier would have to share their Intellectual Property (IP) with Microsoft .
The last condition was as critical as the first, as it would enable the Xbox team to bundle a security system inside the chip (which, unlike the original Xbox, would be properly shielded). Furthermore, the IP rights would grant Microsoft a choice of third-party manufacturers to handle demand and negotiate better costs, resulting in a new console with competitive pricing.
Thus, Microsoft began meeting with Intel, though the talks didn't last long, as Intel wasn't willing to give out its secret recipe. So, Microsoft kept looking.
Resentful old friends

An IBM PC that I found in the Computer History Museum (Mountain View, California), during my visit in June 2019. For some reason, they don't allow you to use it...
It so happened that one of the potential candidates for Microsoft was none other than IBM. Maybe I watched too many dramatic documentaries, but I always pictured the two as the kind of passive-aggressive 'friends' who only smile at each other if they are among other people.
You see, in the past, IBM and Microsoft formed many partnerships that ended up with bitterness between the two. Be that as it may, their agreements ended up disrupting the status quo of the global market and paved the way for new types of consumer products and developments. To name a few examples:
- The arrival of the IBM Personal Computer in 1981 led to a new market of 'personal computers' and the subsequent choice of Microsoft's 'MS-DOS' as the default operating system. This granted Microsoft a monopolistic dominance in the PC business in the years to come.
- In the 90s, with the advent of home networking and affordable servers and workstations, IBM and Microsoft's new venture resulted in OS/2, a new operating system targeting high-end IBM computers. However, upon the unexpected release of Microsoft's Windows NT, IBM was left alone to fight in an arena already favouring Windows.
By the turn of the century, IBM returned its attention to in-house products, including a partnership with Apple and Motorola that resulted in a series of CPUs with a unique architecture called PowerPC.
When Microsoft approached IBM for a new frictionless venture, IBM was already working with Toshiba and Sony to create a superior PowerPC processor for scientific applications. Nevertheless, IBM was open to new business opportunities.
The new CPU partner
In a turn of events, IBM agreed to share its IP and to design a new multi-core processor, and so the Xbox 360's CPU supplier became IBM. Although, you may remember this was the same IBM that already signed an agreement with Sony and Toshiba to produce the PlayStation 3's CPU ('Cell'). Apparently, IBM assumed Microsoft was not aware of the Cell project , and their current contract with Sony did not forbid them from selling to third parties.
All three companies [IBM, Toshiba and Sony]... legally all had rights to go and put any of that technology, any of those processor cores into other spaces. (...) It is very common to develop an interesting, leading-edge new technology and then utilize that technology across multiple platforms. (...) I guess what everyone didn't anticipate was - before we even got done with the Cell chip and PS3 product - that we'd be showing this off specifically to a competitor.
- David Shippy, chief architect of the Power Processing Unit (PPU)
Ironically, as of 2022, IBM's PowerPC chips have vanished from both desktop computers and video consoles, maybe this set a bad precedent and greatly affected IBM's trust in future businesses? I go over this effect later on.
To sum it up, IBM signed an agreement with Sony and Toshiba to develop Cell in 2001. Two years later, in 2003, IBM agreed to supply Microsoft with a new low-powered multi-core CPU. Microsoft's CPU will be called Xenon and will inherit part of Cell's technology, with extra input from Microsoft (focusing on multi-core homogeneous computing and bespoke security). Also, while IBM would brand Cell with its 'BladeCenter' line of servers, Xenon could only be fitted on an Xbox 360 motherboard.
Cell's half-sibling
Now that we've positioned Microsoft and IBM on the map, let's talk about the new CPU. This is how Xenon materialised at the end of the Xbox 360 project...

For comparison purposes, here's the equivalent view of Cell. Cell also includes 32 KB of ROM not shown here.
Don't worry, all of these components will be explained throughout this article, starting with the PowerPC Processor Element (PPE) blocks shown at the top left corner.
A new look at CPU history
The Cell project, with its obsession with vectorised computations, introduced very interesting proposals to radically tackle the ongoing constraints that hindered technological progress. These constraints, however, are very complex. So it wouldn't be fair to say that Cell's methods were the only -or even best- solution. In other words, by studying Xenon's architecture (which competed side-by-side with Cell's) we can gain a different perspective on how CPU architecture evolved throughout that era (the early 00s) and subsequently influenced the next decade of CPUs.

Xenon chip, surrounded by an army of decoupling capacitors.
In doing so, you'll perceive that Xenon takes a more conservative approach than Cell. If we take another look at the previous diagram of Xenon, you can notice that the latter is equipped with the famous PowerPC Processor Elements (PPEs), which is also the most important piece of Cell. However, Xenon's is now equipped with three of them. Additionally, the Synergic Processors Units (SPUs) are no more.
After all, Microsoft didn't want processors of very different natures squashed in their CPU. They instructed IBM to compose three powerful cores and enhance them with the ingredients game developers would expect to find. With this approach, IBM and Microsoft were also able to add non-standard features without disrupting the traditional modus operandi of developers.
Truth be told, this also resulted in aggressive budget cuts to keep this design (and the rest of the system) at a competitive price range. To put it in context, multi-core CPUs for PCs weren't on the store shelves while IBM was building Xenon, and when they debuted in 2005 (coincidentally, the same year the Xbox 360 reached the stores), AMD priced their cheapest Athlon X2 at $537 (equivalent to ~£452 in 2021 money) and Intel charged $241 (equivalent to ~£203 in 2021 money) for their low-end Pentium D - and let's not forget the box only included the CPU.
How this study is organised
We'll now take a look at the main components that comprise Sony's counterpart. To avoid repeating existing information, I'll focus on the novelties of Xenon.
Having said that, the new CPU runs at 3.2 GHz and it includes so much circuitry that, for this study, we have to split it into different groups:
- The three leaders that execute the program's instructions. At first, each resembles Cell's PowerPC Processor Element (PPE), but you'll soon see that they are actually a superset of it. Additionally, since we've got three of them now, it may seem as if the whole chip behaves like a Cerberus monster, where each core may claim control of the whole system. Alas, that's not feasible in a computer, so the first core is the designated master core while the others will be taking assistant roles.
- A single interface that interconnects the cores with the rest of the system. This bus is called XBAR (pronounced 'Crossbar').
- Like in Cell, there are other proprietary interfaces used for debugging or maintenance (i.e. temperature) but these will not be mentioned until we reach the 'Anti-piracy' section.
- The security block which Microsoft oversaw to implement the anti-piracy system. It's a very complex section, so to avoid overwhelming you with information, I'll explain it in the 'Anti-piracy' section as well.
The different approach for Xenon
To explain the aforementioned groups, I've organised the study of Xenon into these areas, in that order:
- The bus connecting all the cores, the XBAR and its special L2 cache block.
- The new refinements of the PowerPC Processor Element (PPE).
- The unusual abundance of general-purpose memory.
- The new programming model suggested (and, in some ways, enforced) by Microsoft.
Inside Xenon: The messenger
The original chip (Cell) was required to house twelve independent nodes actively moving data around, this forced IBM engineers to devise a complicated system that could tackle emerging bottlenecks, which materialised in the form of the Element Interconnect Bus (EIB). With the Xbox 360, Xenon only accommodates three units (the three PPEs), so the EIB has no purpose here. Thus, a simpler solution called XBAR was produced to focus solely on the three PPEs, with space for extra functionality.
XBAR relies on a mesh topology that doesn't direct traffic in a ring-style manner. Instead, each node is provided with a dedicated lane to move its data through . This may appear more optimal than the ring topology of the EIB, but that's because the XBAR only needs to serve a small number of nodes. Furthermore, the XBAR operates at full speed (3.2 GHz).

Simplified diagram of the XBAR/Crossbar combined with the L2 component.

For comparison purposes, this is the architecture of Cell's EIB (found in the PS3).
To be fair, until now I only talked about the particular interface that interconnects the PPEs. Well, the XBAR is just one piece of the sizeable chunk IBM designed for Xenon. It turns out the leftover space gave them room to incorporate another very important block that speeds up the transactions between the PPEs and the rest of the system: L2 Cache.
Shared cache
Once the PPEs pass through the XBAR, they get access to 1 MB worth of L2 cache. The L2 block is also connected in a meshed manner, where each PPE gets a separate 256-bit bus but is now clocked at 1.6 GHz (half the PPEs' speed) .
Coincidentally, Cell also houses 512 KB of L2 cache for its single PPE. I only described it in one sentence in my previous article. Now, however, we find ourselves with a larger space that's also shared between three cores, so I find it necessary to dive into the properties of the L2 block so we can understand how it will condition the performance of the PPEs.
First things first, every time the cache fetches data from memory (in the event of a 'cache miss'), it does so by pulling a large slice called 'cache line', which is 128 Bytes wide in the case of Cell and Xenon. Then, L2 records the cache line on an internal list for locating it in the future. Moreover, in Xenon/Cell, L2 is 8-way associative, which means that cache sets may store up to eight different cache lines. Don't worry if you don't know what this means, the theory behind CPU cache can be hard to follow, especially if you only want to learn about game consoles. In layman's terms, the greater the number of associations, the fewer probabilities of cache misses, but the slower it becomes to iterate through the internal list.
The choice of an 8-way associative cache was not a rash decision for Xenon, as providing eight associations can alleviate the six simultaneous threads (each PPE is dual-threaded) trying to access the L2 block at the same time. This also balances frequent cache misses and lookup times. All of this, while keeping the costs down. For comparison purposes, the expensive Intel 'Smithfield' (a Pentium D from 2005) provides two cores with 2 MB of L2 cache each!
It wouldn't be the first time Microsoft trimmed the cache for pricing purposes, so it will be up to developers to optimise its usage. To assist with this, XBAR bundles additional logic like 'cache locking' which will be further explained in the 'Graphics' section.
Inside Xenon: The leader(s)
This is as far as we go with our description of the XBAR and L2 cache, let's now talk about the actual CPU. Just a reminder that, since I've already dissected the standard PPE (found in Cell) for the PS3 article, I'll be focusing on the novelties here.
To start with, Xenon's PPEs don't feature a PowerPC Processor Storage Subsystem (PPSS) anymore, presumably since the interfacing part is handled by the XBAR and the L2 cache is now shared across the three units.

Simplified diagram of Xenon's PowerPC Processor Element (PPE).

For comparison purposes, this is Cell's PPE.
To be honest, I'm not sure why technical manuals keep calling Xenon's PPEs a 'PPE' as they better resemble a PPU.
Anyway, let's go over the significant changes that Microsoft liked to brag about when being compared with the PlayStation 3.
The new vector units
As the Synergistic Processor Elements (SPEs) were Sony's secret weapon and Microsoft was only interested in homogeneous systems, IBM presented an alternative approach to speed up the manipulation of vectors and matrices in Xenon . In a nutshell, IBM boosted the VMX unit (the SIMD block found within the PPEs) which evolved into VMX128, now housing more registers and opcodes.
The initial VMX specification implemented in the PlayStation 3 provides 32 128-bit registers and instructions for operating up to three 32-bit scalars. This worked to an acceptable degree with general-purpose applications that depend on SIMD operations, although the real performance would be unlocked once the SPEs are added into the equation. By contrast, Microsoft wanted programmers to port SIMD-hungry applications without extra hassle, and the new VMX128 unit reflects that.
VMX128 supplies 128 128-bit registers instead, along with an adapted instruction set to manipulate the larger set of registers. To accomplish that, IBM changed the opcode format to allocate 7 bits (as opposed to 5 bits) for referencing its extended register file . This is possible thanks to some trickery applied on the last five bits at the end of the 32-bit opcode . The last five bits are mostly, but not completely, unused by VMX. Consequently, VMX128 is incompatible with a subset of VMX instructions (related to integer multiplication and additions).
Furthermore, VMX128 adds new instructions that compute the dot product of two vectors composed of up to three 32-bit floating-point numbers; and others that handle Direct3D's data compression formats (it's worth mentioning that DirectX is the sole Application Programming Interface (API) for programming this console). Finally, thanks to the symmetric design of Xenon and its multi-threaded model (dual-issuing), VMX128's register file is duplicated, so there are 256 128-bit registers per core!
As we reach the end of this section, there's still one question left unanswered: which is faster for vector operations, 1 VMX + 6 SPEs (as in the PS3) or 3 VMX128 units (as in the Xbox 360)? Well, their designs are too divergent, so it's hard to quantify. One could say 'The SPE can execute up to two instructions per cycle while Xenon takes 12 cycles to add two vectors (due to the long pipeline of the PPE)' but that's relative, as the SPE's memory scope is restricted to its local memory (requiring DMA calls to interact with the outside), while Xenon's PPEs can access any memory location. So, in conclusion, these are two contrasting models and programmers will just have to get the best out of them.
A new but short-lived instruction
With the advent of a larger cache system, Microsoft and IBM extended the PowerPC instruction set to accommodate some instructions that operate these blocks from the program side, in case programmers wish to make manual interventions (such as caching data ahead of time or saving L2 space). In any case, the standard PowerPC specification already provides the dcbt ('Data Cache Block Touch') instruction to fetch data from memory into the L1 cache (individual for each core) . Nonetheless, Microsoft took a step forward and added xdcbt ('Extended Data Cache Block Touch') to fill the L1 cache without even going through L2 (shared across all cores) . This makes sense as the shared L2 is a new addition of Xenon. However, bypassing L2 is a risky operation as two cores may end up seeing different data (cache incoherency), so it will require correct handling from the program side to keep cache incoherency and race conditions out of the way.
As luck would have it, xdcbt worked fine until Microsoft started receiving error reports from game studios . Initially, Microsoft relied on this instruction for common routines available through their API (such as memcpy()) but they didn't account for the consequent lack of cache coherence. Thus, uses of xdcbt were subsequently removed from their APIs. Yet, studios found a bigger problem: branching instructions followed up by xdcbt always execute the latter. It turns out the branch predictor will attempt to execute xdcbt (as any other instruction that may be undone later on, if the prediction turns out to be incorrect) with the exception that once the cache blocks are mangled, there's no way to undo it (becoming a hazard). Thus, the program will not synchronise the caches as it assumes xdcbt hasn't been triggered, leaving the PPEs in a disarrayed state.
In the end, Microsoft purged xdcbt from the compiler due to its non-deterministic (thus unpredictable) behaviour, and what's left of it is just an anecdote.
Revisiting old paradigms
There's a recurring subject found in noteworthy writings from 2005 like 'Inside the Xbox 360' by Jon Stokes or 'Understanding the Cell Microprocessor' by Anand Lal Shimpi , and that is the lack of out-of-order execution that once debuted in early PowerPC chips like Gekko, but for some reason is completely absent in Cell & Xenon. If you recall from the GameCube article (which I wrote two years ago), Gekko is an out-of-order CPU, meaning it's able to analyse the instruction stream as instructions come in, and subsequently reorder them to better distribute the load on Gekko's internal units.
By then, CPU cores employing out-of-order execution were in the order of the day (pun intended). IBM's PowerPC 604 (1994) brought it to high-end Macintosh computers, Intel's P6 (1995) introduced it to the x86 line and MIPS implemented it with the R10000 (1996) CPU, a successor of the R4000 (found on the Nintendo 64). Afterwards, all of a sudden, Cell and Xenon arrive with an in-order execution style... care to explain?
To understand this turn of events, let's take a closer look at the Out-of-order execution (OoO) model. OoO is not a Lego piece that engineers fit and then move on to the next part. Pipelined CPUs are susceptible to a whole new range of hazards, which means that the slightest addition may lead to a complete redesign. For instance, OoO requires altering the register file to fit register renaming , a clever technique where the CPU houses more registers than the instruction set references, enabling the CPU to store multiple copies of the program state and prevent dependency problems that arise after altering the original order of the instructions. All in all, while these solutions are interesting, they add complexity to the CPU chip, and that can pose a threat to future scalability.
The PowerPC Processor Element (PPE) derives from POWER4 which is itself an OoO core. However, POWER4 is an independent (and power-hungry) chip by itself, while the PPE is either surrounded by 7 SPEs (as with the PlayStation 3/Cell) or by two more PPEs (as with the Xbox 360/Xenon). These changes take up space and the dimensions of the CPU chip are critical in terms of price and heat emission. So, in the end, I imagine IBM didn't see enough advantages for applying OoO on an application-specific machine like the PS3 or Xbox 360.
An alternative solution
Look at it this way, why was OoO invented in the first place? To prevent the CPU from idling (as idling = wasted resources). So, will the applications (3D gaming) be prone to this issue? Maybe, but it's also true that OoO is not the only possible solution to this problem! Introducing Thread level parallelism (TLP).
What if instead of sorting out which instructions should be executed first, we duplicate the resources into two (or more) groups called 'threads' and let the program switch between the different groups of resources (multi-threading) as it sees more fit. This is what IBM's engineers went for, hence the reason the PPE is dual-issued. The new technique, called Thread level parallelism (TLP), differentiates itself from Instruction level parallelism (ILP) by letting the program, as opposed to the CPU, to come up with its own solution. With Xenon and Cell, it's now left to the compiler and the program's multi-threading implementation to produce an efficient sequence of instructions.
The interesting thing is that neither approach is better or worse, out-of-order processors are still found in the market (Intel/AMD still supports OoO along with an obscene amount of other techniques, while ARM adopted OoO with the Cortex-A9 in 2007). On top of that, those CPUs have multi-threaded cores and even bundle multiple cores within the same chip, so you get a mix of both techniques (TLP and ILP).
Eroded relationships
Conversely, whilst the omission of out-of-order may look justified from an engineering standpoint, this would ultimately wither trust in IBM's long-term abilities. Apple, sensing that future PowerPC chips would no longer stand against Intel , was the first to abandon its partnership and switch to Intel in 2005 (just in time for Yonah). Nonetheless, this didn't worry IBM, who preferred to focus on its server and gaming sector.
The long-term outcome of the PowerPC CPU, however, became extinction. The aforementioned technical constraints (after all, IBM used to be a trailblazer in out-of-order design), combined with questionable business ethics (the fact IBM secretly designed for two close rivals at once), resulted in the loss of the entire console market: Sony and Microsoft selected AMD CPUs for their next-generation console. A decade later, Nintendo's Wii U would become the last console to bundle a PowerPC chip.
Inside Xenon: Main Memory
Unlike its competitor equipped with 256 MB of XDR DRAM, there's 0 MB of external memory installed next to Xenon. That's right, a 100% decrease compared to the Japanese counterpart.
Okay... let me rephrase that. On the motherboard, multiple chips provide a total of 512 MB of GDDR3 SDRAM, but they are sitting next to the GPU, not the CPU. So... you probably guessed it. Just like the original Xbox, Microsoft went for the Unified Memory Architecture (UMA) layout, where all components share the same RAM chips. This provides more flexibility in the amount of memory reserved for the CPU or GPU. However, a single point of access for various components means the far away components (like the CPU) will suffer higher latency compared to the hypothetical case where the CPU was given dedicated memory. This was duly noted by IBM and Microsoft engineers and subsequently tackled with the aforementioned L2 cache along with extra circuitry (which is explained in the 'Graphics' section).

Xenon next to 'Xenos' (the GPU) guarding 256 MB of GDDR3, the remaining half is on the back.
The type of memory chips (GDDR3) is the same one found in the Wii and PS3. Conversely, the Xbox 360 gets the crown for including the largest amount. In fact, it's been reported that Microsoft initially planned to house 256 MB, but later doubled the amount out of fear for Sony's upcoming competitor . Though, unlike Sony, Microsoft was extra sensible with the retail price, so it had to cut down on other features. That's why the hard drive was made optional. Furthermore, Microsoft expected Samsung (its memory supplier) to deliver their chips operating at 1.4 GHz but then settled for 700 MHz at the time of shipping the console.
Memory Controller
With all being said, how can Xenon access this memory? Well, the CPU communicates with the GPU through an interface called Front-side Bus, this combines serial and parallel communication models to keep costs down and reduce latency as much as possible .
On the outside, there are two unidirectional lanes named 'PHY', each is made of 16 buses that operate at the exceptional speed of 5.4 GHz . In there, information is transmitted in serial form. Internally, however, the CPU and GPU only understand whole words. Thus, both endpoints are tasked with serialising the data before sending it through PHY and/or deserialising it after receiving it. The CPU's inner interface is 64 bits wide and runs at 1.35 GHz, while the GPU's is 128-bits wide and runs at 675 MHz . If we do the math, both lanes provide a bandwidth of 10.8 GB/s.
The Front-side bus route only takes you to the GPU. So, between the GPU and the GDDR3 chips, two memory controllers inside the GPU manage this connection using one 1024-bit bus each. To reduce latency, the memory controllers use address tiling for GPU-related operations and path-finding to reduce CPU congestion. In general terms, Microsoft states that there's a bandwidth of 22.4 GB/s between GPU and memory.
It's all jolly on paper, but let's not forget that the CPU still has to walk a long road to get to memory. To give you an idea, if you grab PIX (the CPU profiler included with the SDK) and run the sample tests that come with the utility, you'll see that for every cache miss, the CPU spends ~600 cycles to get to memory! This is one of the features of UMA-based systems you don't see marketed very often :), but also helps you understand why cache is so critical.
Inside Xenon: Programming styles
With all being said, how can anybody take advantage of this fascinating chip? Well, there's only one way: Multi-threading.
From the programming perspective, a CPU that's made of multiple homogeneous cores sharing the same memory is what we call a Symmetric Multi-Processing (SMP) design. This conditions how the application will be designed and implemented. As of 2022, this is the de-facto layout for consumer multi-core CPUs (i.e. x86 and ARM) as well.
SMP programming abstracts access to physical CPU cores with the use of 'virtual threads'. A virtual thread is a sequence of instructions the programmer defines. Threads can then be submitted to a 'scheduler' for their execution. The scheduler is another program (often part of the operating system) that handles how virtual threads are dispatched to the physical CPU cores.
This abstraction layer allows the programmer to avoid discriminating against the type of core used (therefore making the program cross-compatible with similar platforms) and hard-coding the number of cores (making it scalable).

Representation of the multi-threading paradigm on Xenon. A program may create n-threads (two in this example), and then the operating system's scheduler takes care of dispatching the threads to physical cores. Bear in mind, the operating system also runs as a thread.
It's no surprise that this style became a standard in future console generations, as it's easier to scale and program with symmetric systems than asymmetric ones (i.e. Cell). The latter often depends on unusual programming styles, eroding compatibility with other systems.
Also, it's worth mentioning that, while SMP programming is not necessarily restricted to one platform, Xbox 360 developers will only be writing threads for Xenon (which provides a total of six threads). This enables them to optimise their multi-threading design for the Xbox platform as well.

