Follow one 3D point across various situations. You can think of it as the journey of an Asset from just being an Asset to how the 3D asset is represented on a 2D screen. Along the way we'll derive the linear algebra that works behind all these transformations.
What this lesson assumes from Part I
Vectors, unit vectors, the dot product, and the cross product.
Affine 4×4 transforms with a fourth coordinate w = 1, including translation, rotation, and scaling.
Matrix composition is right-to-left: AB applies B first. Inverse transforms undo a placement.
Coordinate spaces, if dumbed down, means just point of view. You can see a bird sitting on a branch from the ground, or from a building across the street, your view changes and so does the bird's coordinates, but the bird itself is the same. A coordinate space is a choice of origin and axes used to describe a point. The physical point does not change when its coordinates are rewritten in another space.
The three spaces
Model space belongs to one object. The cube’s own origin and axes are used. This is the actual Asset as it exists and captured in a 3D model, waiting to be placed in the virtual world.
World space is the shared room containing every object and the camera. This is the virtual world where the assets are placed, and they interact with each other and the camera, telling the story of the scene.
Camera space, also called view space, uses the camera as the origin. This is the point of view of the camera. It's like you are sitting in the camera and looking out, seeing the world from that perspective. It's x-axis points right, its y-axis points up, and it looks along −z.
Let's try to understand it using the playground below. P is one physical point: the labeled orange corner of the cube. A 4×4 matrix rewrites P into the next space. Mmodel places the object in the world; Mview rewrites world coordinates relative to the camera.
Pworld=MmodelPmodelPcamera=MviewPworld
The camera’s pose is where it sits in the world and which way it faces (the black pin in the World pane). If you used that pose as a matrix, it would plant the camera into the world.
Mview does the opposite: it leaves the camera at the origin and moves the cube instead. In the World pane, the pin moves around a fixed cube. That's because you are looking at the world from the camera's perspective, so the cube appears to be moving around the camera. In the Camera pane, the cube moves around a fixed pin.
This section uses Mview for its job. The next section constructs every row of that matrix from the camera’s position and axes.
P (orange) is a point: one corner of the cube, at (1, 1, 1) in model space. Each matrix M below rewrites that same point.
The teal cube is the object. It is the same cube in every frame, just written in different coordinates.
The black pin is the camera’s position (the eye). The camera icon faces along its −z-axis.
The dashed gray cube in World is a ghost of Model: where the cube sat before Mmodel.
Walkthrough: one corner, three names
Model. The cube sits on its own origin. P is just (1, 1, 1). There is no camera in this space, so this pane never changes when you drag the sliders.
World.Mmodel placed that cube in a shared room. The orange dot is the same corner, now in world coordinates. The black pin is the camera standing somewhere in that room. The dashed orange line is “camera to P.”
Camera. Undoing the camera’s placement locks the pin at (0, 0, 0), looking along −z, while the cube moves. Drag camera x: in World the pin slides, while in Camera the cube slides the other way. It is the same motion described from the opposite frame. That is Mview.
Playground: all three frames at once
1 · Model
P = (1.00, 1.00, 1.00)
Local coordinates. No camera. Sliders do not affect this pane.
2 · World
P = (1.70, 1.48, -0.24)
Shared room: cube after Mmodel, camera as a black pin. Drag the camera; P stays put.
3 · Camera
P = (1.59, 2.92, -3.52)
The pin is locked at (0, 0, 0), looking along −z. Drag the sliders: the cube moves, not the pin.
camera x
slidernumber input
camera z
slidernumber input
yaw (turn left/right)
slidernumber input
Move camera x and watch the World and Camera panes together.
Build the camera’s coordinate frame
The previous section used a camera matrix. We can now build it using only vector subtraction, normalization, cross products, and dot products from Part I.
Three inputs
eye is the camera’s position in world space.
target is a world-space point the camera looks toward.
up is a preferred upward direction. Here it is the world y-axis direction (0, 1, 0).
Step 1 · build three perpendicular unit axes
We name the camera’s right, up, and backward axes u, v, and w. The camera looks along −w.
w=normalize(eye−target)
eye - targeteye − target points backward, from the target to the camera. Crossing the preferred up direction with w gives a right-pointing vector. Normalize it to make its length 1. A final cross product gives the corrected camera-up axis.
u=normalize(up×w)v=w×u
Step 2 · find a point’s camera coordinates
Let p be any point in world space. First subtract eye so the camera position becomes the origin. A dot product with each camera axis then measures the point along that axis:
xc=u⋅(p−eye)yc=v⋅(p−eye)zc=w⋅(p−eye)
Distribute each dot product. For example, xc = u·p − u·eye. The coefficients of p become the first three entries of a matrix row; the remaining number is the translation entry.
Camera space left the pin at the origin, looking down −z. The rewritten scene still includes points behind the camera, off to the side, and infinitely far away. A picture cannot include all of them. The view volume is the region we keep.
The first boundary is the near plane. Its rectangle is the film: the window through which the scene is viewed. l, r, b, t name its left, right, bottom, and top coordinates. near and far are positive distances from the camera. Because the camera looks down −z, the two planes have camera-space coordinates z = −near and z = −far.
The six numbers define two possible visible regions. For perspective, the side walls open away from the camera and form a frustum, meaning a pyramid with its tip cut off by the near plane. For orthographic projection, the walls stay parallel and form a box.
A point outside the chosen volume is rejected before drawing. This rejection is called clipping. A point behind the camera is also rejected because no forward-looking ray reaches the film.
What you are looking at
Both panes are still camera space. The black pin is the same origin as the Camera pane above.
The teal quad is the near plane (the film). The orange quad is the far plane.
Dots: teal = inside that volume, gray = outside (rejected), and orange = behind the camera. Watch B: it can sit in the wedge and miss the box.
Why this exists
ApplyMview. We then know where every point is relative to the camera, but we still do not know which points fall within a finite image.
Keep a window. Only points that pass through the near rectangle, and lie between near and far, are drawn. The rest are thrown away (clipped).
Two visible regions, one window. Both shapes share the same six numbers. A photograph wants the wedge because objects appear smaller as their distance from the pin increases. A CAD drawing uses the box because size does not change with depth. Drag r and both volumes widen together.
Playground: wedge and box at once
Perspective
Walls open from the pin. The far end is larger. B often fits here even when it misses the box.
Orthographic
Parallel walls. The far end is the same size as the film. B is often outside this box.
l
slidernumber input
r
slidernumber input
b
slidernumber input
t
slidernumber input
near
slidernumber input
far
slidernumber input
A(ahead):perspectiveinside·orthographicinside
B(to the right):perspectiveinside·orthographicoutside
Look at B in both panes, then drag r until the box contains B too.
One standard target: the canonical cube
The two view volumes have different shapes, positions, and sizes. Later drawing steps would be needlessly different if they had to understand every possible camera volume. Instead, each projection will rewrite its chosen volume into the same standard cube.
That standard target is the canonical view volume. Every coordinate in it lies between −1 and +1. Coordinates measured in this cube are called Normalized Device Coordinates, abbreviated as NDC.
−1≤xndc≤1−1≤yndc≤1−1≤zndc≤1
The six boundary rules
Horizontal. Left becomes x = −1; right becomes x = +1.
Vertical. Bottom becomes y = −1; top becomes y = +1.
Depth. Near becomes z = +1; far becomes z = −1.
These endpoint rules are the facts from which we will derive every matrix entry. No projection matrix has been used in the diagram below.
Diagram: starting volume and required NDC target
Camera (view) space
The wire shape is the volume selected in the previous section. Teal and orange are equal-size example objects, not part of the volume itself.
Required target: NDC
This is the destination, not a result computed early. The next sections derive the two maps that reach it.
Switch only the left starting shape. The labeled NDC target on the right stays the same.
Orthographic projection
Orthographic projection preserves x and y without shrinking them according to depth. A closer cube and a farther cube are therefore rendered at the same size. Mapping the view box onto [−1, 1]³ uses one scale and one translation per axis. The axes are not mixed: output x uses only input x, output y uses only input y, and output z uses only input z.
To shorten the matrix, let n = −near and f = −far. They are signed z coordinates, not positive distances. Because the camera looks down −z, both are negative and n > f.
n=−nearf=−farn>f
One axis, then copy it three times
Take a number u that ranges from u₀ to u₁. We want a new number U that ranges linearly from −1 to +1. Shift by the midpoint so the interval is centered on 0, then divide by the half-width so the ends land on ±1.
What the symbols mean
u is the original x, y, or z coordinate.
u₀ and u₁ are the lower and upper endpoints of the input interval.
m is the interval midpoint; h is its half-width.
U is the normalized output coordinate, between −1 and +1.
s is the scale factor; c is the constant translation after scaling.
We require u₁ ≠ u₀. If the endpoints are equal, the interval has zero width and cannot be stretched onto [−1, 1].
Averaging the endpoints finds the center. Half of their difference finds the distance from that center to either endpoint.
m=2u0+u1h=2u1−u0U=hu−m
Subtracting m moves the midpoint to 0; dividing by h makes each half of the interval one unit long. Thus u₀ maps to −1, m maps to 0, and u₁ maps to +1.
u=u0⇒U=−1u=m⇒U=0u=u1⇒U=1
Rewrite that as U = s u + c, which is exactly one row of a 4×4 matrix:
The scale s changes the interval’s size; the translation c moves its center to zero. The zeros in the matrix say x does not use y or z, and so on.
Worked example: a centered interval
For u₀ = −1.20 and u₁ = 1.20, the midpoint is 0 and the half-width is 1.20:
m=2−1.20+1.20=0h=21.20−(−1.20)=1.20
s=2.402=65c=0U=65u
Input u
Calculation
Output U
−1.20
(5/6)(−1.20)
−1
−0.60
(5/6)(−0.60)
−0.5
0
(5/6)(0)
0
0.60
(5/6)(0.60)
0.5
1.20
(5/6)(1.20)
1
Because this interval is already centered at zero, no translation is needed: c = 0.
Worked example: translation required
For u₀ = 2 and u₁ = 6, the midpoint is 4 and the half-width is 2:
m=22+6=4h=26−2=2U=2u−4=21u−2
Here s = 1/2 and c = −2. Substitution verifies the whole map: 2 → −1, 4 → 0, and 6 → +1.
u=2⇒U=−1u=4⇒U=0u=6⇒U=1
x, row 1. u₀ = l, u₁ = r. Scale 2/(r − l), translation −(r + l)/(r − l). Left wall → −1, right wall → +1.
y, row 2. u₀ = b, u₁ = t. Scale 2/(t − b), translation −(t + b)/(t − b). Bottom → −1, top → +1.
z, row 3. Near should become +1 and far should become −1. The near-plane coordinate n might be −2, while the far-plane coordinate f might be −8, so n > f. Apply the same formula with u₀ = f and u₁ = n: scale 2/(n − f), translation −(n + f)/(n − f).
Fourth-coordinate row. A 4×4 matrix carries one extra coordinate, initialized to 1 for a point. The row (0, 0, 0, 1) leaves it equal to 1. Perspective will introduce and use this coordinate later.
How the four rows are earned
matrix row
required output
entries
1
x NDC from x
l → −1 and r → +1
[2/(r−l), 0, 0, −(r+l)/(r−l)]
2
y NDC from y
b → −1 and t → +1
[0, 2/(t−b), 0, −(t+b)/(t−b)]
3
z NDC from z
n → +1 and f → −1
[0, 0, 2/(n−f), −(n+f)/(n−f)]
4
keep w equal to 1
ordinary points enter with w = 1
[0, 0, 0, 1]
z=n⇒n−f2n−(n+f)=1z=f⇒n−f2f−(n+f)=−1
Assembling the four rows gives the orthographic matrix. Every entry above is now accounted for.
The square on the right is the canonical cube viewed along the camera’s −z direction: x points right and y points up. The depth coordinate z is stored too, but is not drawn as a third axis in this square.
The 3D pane is an observer standing beside the volume, so near and far sit left and right. The NDC pane is what the pin sees. The two panes show different views. Under orthographic projection, x and y ignore z, so the front and back of a cube land on one square. Teal (near) and orange (far) are the same size; they only shift if their x or y coordinates differ.
Playground: orthographic box, NDC, and a test point
Perspective makes distant objects look smaller. We will first derive the geometric rule, then show how a 4×4 matrix and one division reproduce it.
Symbols in the diagram
P = (x, y, z) is the original point in camera space.
P′ = (x′, y′, n) is where the ray from the camera through P meets the near plane.
x and y locate P horizontally and vertically. z is its signed depth and is negative in front of the camera.
n = −near is the signed z coordinate of the near plane, introduced in the orthographic section.
Similar-triangle derivation
The camera is at the origin. P and P′ lie on the same straight ray. In the x–z side view, the small triangle has side lengths |x′| and |n|; the larger one has matching lengths |x| and |z|. Similar triangles have equal ratios:
∣n∣∣x′∣=∣z∣∣x∣
Lengths alone do not say whether P is left or right of the center. Restoring the signed x coordinates gives x′/n = x/z. Multiply both sides by n:
x′=znx
Looking from above gives the same argument for y. Therefore, the point on the near plane is:
P′=(znx,zny,n)
For a visible point, n and z are both negative, so n/z is positive. As |z| grows, n/z gets smaller and the image moves toward the film’s center.
Playground: drag P and watch the similar triangles
P horizontal x
slidernumber input
P depth z
slidernumber input
near distance
slidernumber input
far distance
slidernumber input
P = (1.10, 0, -4.00)
x′ = n·x/z = 0.55
Why a fourth coordinate w is useful
Part I represented an ordinary 3D point as (x, y, z, 1). The fourth number is called w. It is a homogeneous scale coordinate, not another spatial direction. Part I kept w equal to 1; perspective deliberately changes it.
A four-number homogeneous point becomes an ordinary three-number point by dividing its first three numbers by w. This operation is the perspective divide:
(xh,yh,zh,w)⟶(wxh,wyh,wzh),w=0
Multiplying all four numbers by the same nonzero α does not change the three ratios. The two four-number tuples below therefore represent the same ordinary point:
Call the matrix Mwarp. Its input is the column vector (x, y, z, 1). We choose each row by stating exactly what its output must do.
Rows 1, 2, and 4 · create the perspective fractions
We need x′ = nx/z after the divide. Make the first output nx and make the fourth output w = z. Dividing the first by the fourth then gives nx/z. The same choice gives ny/z vertically.
xh=nxyh=nyw=z⟹wxh=znx,wyh=zny
Row 3 · keep the near and far depths fixed
Let the third row produce zh = Az + B. After division by w = z, the ordinary depth is A + B/z. We require the near input z = n to remain n and the far input z = f to remain f:
A+nB=nA+fB=f
Subtract the second equation from the first. Solving gives B = −fn. Substituting that into either equation gives A = n + f.
B=−fnA=n+fzh=(n+f)z−fn
Every row follows one stated requirement
matrix row
required output
entries
1
xₕ = nx
xₕ/w must become nx/z
[n, 0, 0, 0]
2
yₕ = ny
yₕ/w must become ny/z
[0, n, 0, 0]
3
zₕ = (n+f)z − fn
near remains n; far remains f after ÷w
[0, 0, n+f, −fn]
4
w = z
the divide must use the point’s depth
[0, 0, 1, 0]
Mwarp · derived rows · read each row from left to right
Mwarp changes the frustum into an orthographic box while keeping the near and far planes in place. The already-derived Morth then maps that box to the NDC cube. Part I’s right-to-left composition rule says the combined matrix is:
Mper=MorthMwarp
Multiplying the two matrices row by column gives the full perspective matrix below. Its four-number output, before dividing by w, is called the point’s clip coordinates. Dividing clip coordinates by w gives NDC.
Mper · full perspective projection · read each row from left to right
Playground: α changes clip numbers, not the NDC point
homogeneous scale α
slidernumber input
input point = (1.10, 0, -4.00, 1)
clip = (-1.83, 0.00, 1.33, -4.00)
α-scaled clip = (-2.93, 0.00, 2.13, -6.40)
after ÷w: (0.46, 0.00, -0.33) = (0.46, 0.00, -0.33)
The default α is 1.6. Every clip coordinate becomes 1.6 times larger, including w, so each ratio after ÷w stays unchanged.
Why perspective depth is not evenly spaced
Row 3 of Mwarp, followed by the divide by w = z, produces (n + f) − fn/z. Applying Morth afterward sends near to +1 and far to −1. The 1/z term means equal changes in camera-space depth do not create equal changes in NDC depth.
zafterwarp=(n+f)−zfnz=n⇒nz=f⇒f
Diagram: nonlinear depth after the full perspective projection
near z = -2.00 → NDC +1 · far z = -8.00 → NDC −1 · current z = -4.00 → NDC -0.33
Orthographic vs perspective
The same near and far cubes are shown through two cameras. Orthographic projection preserves their screen size. Perspective projection makes the farther cube shrink because the rays meet at the camera origin.
Both matrices have now been derived. This playground applies them to identical geometry so only the projection rule changes.
Green is nearer, and orange is farther. Switch modes and watch the projected output.
One point through the complete pipeline
This section places the operations in sequence so we can follow a point through every named stage.
A viewport is the rectangular screen region where the final image is drawn. Its width W and height H are measured in pixels. The last stage maps the NDC square into this rectangle.
The six stages
Model. P starts in the object’s own coordinate space.
World. Mmodel places the object in the shared scene.
Camera (view). Mview rewrites P relative to the camera.
Clip. Mper produces four clip coordinates. They have not yet been divided by w.
NDC. Divide clip x, y, and z by clip w. The visible canonical cube is [−1, 1]³.
Screen. The viewport map changes NDC x and y into pixel coordinates.
Screen x increases to the right, but screen y usually increases downward.
For x, require NDC −1 → pixel 0 and NDC +1 → pixel W. The same linear endpoint calculation used for the orthographic matrix gives scale W/2 and translation W/2.
xscreen=2Wxndc+2W=2(xndc+1)W
For y, require NDC +1 at the top to become pixel 0, and NDC −1 at the bottom to become pixel H. The negative scale performs the direction flip.