Partial derivatives and the gradient vector extend differentiation to functions of several variables, and with them arrives the geometry of surfaces, tangent planes and directional rates that dominates applied mathematics from mechanics to machine learning. A partial derivative is a one-variable limit taken with all the other coordinates held fixed, so a differentiable function of three variables has three of them at a point, each answering the question of how the function responds when one coordinate moves alone. The gradient collects these quantities into a vector whose remarkable property is that it points in the direction of fastest increase. For any unit direction the directional derivative, the rate of change of the function as the input moves along a ray in that direction, equals the dot product of the gradient with the direction vector, an expression that attains its maximum exactly when the two vectors are parallel. Temperature fields, gravitational potential, pressure, and the value of a loss function all obey this rule, which is why the gradient appears in thermodynamics, in the equations of elasticity, in fluid mechanics, and at the centre of the training loop of a neural network. The differential of a function of several variables is the linear map that approximates it best near a point, and total differentiability means that the error of that approximation is smaller than the size of the displacement. A function can possess all its partial derivatives at a point and still fail to be differentiable there, so the existence of partial derivatives is a weak guarantee, while continuity of the first derivatives on a region is enough to secure differentiability throughout it, and this is the hypothesis most theorems adopt to avoid technicalities. Where differentiability holds, the tangent plane at a point has the gradient as its normal, and the equation of the plane is written by taking the dot product of the displacement with the gradient, which turns local information about a surface into a usable linear model. Higher partial derivatives form the Hessian matrix, symmetric when the mixed partials are continuous, whose eigenvalues describe the local shape of the surface and whose determinant distinguishes maxima, minima and saddle points in the second derivative test. The multivariable chain rule propagates gradients through compositions, and reverse-mode automatic differentiation is simply that rule applied backwards through a computation graph, accumulating contributions rather than recomputing them, which is the algebraic content of the backpropagation algorithm. Constraints are where local analysis becomes global design. The implicit function theorem says that a smooth relation between variables allows one variable to be solved for as a smooth function of the others wherever a suitable derivative is nonzero, and it underlies the tangent plane approximation to a solution surface, to a reaction rate law, and to a design constraint. Lagrange multipliers then handle the restricted problem: to extremise a function subject to a smooth constraint, one adds a multiplier and solves the enlarged system in which the gradient of the function is a multiple of the gradient of the constraint, so that the level sets of the two touch at the optimum. In engineering this is how a truss of fixed volume is shaped to carry a load with minimum weight, how a chemical equilibrium is located, and how a control trajectory is chosen; in statistics the same idea yields the score equations of a likelihood maximisation; in machine learning it becomes the projected gradient step used in matrix factorisation and in constrained embeddings. Numerical work in this area requires care, because the steepest descent direction is known exactly but the distance one may safely travel along it is not, and this tension between speed and stability reappears in numerical integration and again in the study of dynamical systems. For a student in engineering or computing the payoff is immediate: a finite element stiffness matrix is an assembled Hessian, a gradient descent optimiser is a discrete dynamical system, and a Newton step is a local linearisation, so the theory of partial derivatives is not an ornamental extension of the single variable calculus but the working language of the subject. Three habits sharpen the practice of this subject. Before differentiating anything, decide explicitly which variables are held fixed, since the vague instruction to treat a quantity as a constant is a frequent source of silent error, and check the units, which catches a derivative whose dimensions cannot possibly be right. When a gradient appears inside a physical model, ask what the divergence and the curl of the underlying field mean, because that question usually reveals whether the model was written correctly in the first place. It is also worth remembering that the Hessian describes local behaviour only, and that convexity of a function on a convex set is precisely the condition guaranteeing that any stationary point found by gradient descent is a global minimum, a fact that explains why many optimisers used in practice are restricted to convex objectives or are given safeguards that force them to behave as though the loss were convex.