Reference
XPath 1.0
40 entries covering the functions, axes and operators libxml2 implements — the engine behind PHP's DOM, Python's lxml and Ruby's Nokogiri.
Every example here was run, not written
The result printed on each page is what the evaluator actually returned, asserted on every test run. That matters more than it sounds: an XPath reference whose examples are subtly wrong costs an hour before you stop believing it. Run your own.Functions · 21
- count()count(node-set) → numberHow many nodes a node-set contains. The fastest way to check whether an expression matches what you think it does.
- last()last() → numberThe size of the current context, which makes [last()] the idiom for the final node in a set. There is no first(): use [1], because [0] selects nothing.
- position()position() → numberThe one-based index of the context node, so [n] is shorthand for [position() = n]. [1] is the first in each context, not the first in the document.
- local-name()local-name(node-set?) → stringAn element's name without its namespace prefix — and the standard way to match namespaced documents when you cannot bind a prefix.
- name()name(node-set?) → stringThe node's qualified name, prefix included — which makes it dependent on how the document happens to be written.
- namespace-uri()namespace-uri(node-set?) → stringThe namespace URI a node is in — the half of its identity that carries meaning. Compared as an exact string, so a trailing slash is a different namespace.
- id()id(object) → node-setSelects elements by their DTD-declared ID, which means it does nothing in most documents: an attribute merely named 'id' is not an ID to XPath.
- contains()contains(haystack, needle) → booleanWhether one string occurs inside another. The most-used XPath function, and the one most often given its arguments the wrong way round.
- starts-with()starts-with(haystack, prefix) → booleanWhether a string begins with another, tested string-first. XPath 1.0 has no ends-with() — that is 2.0, and it reports as an unregistered function.
- substring()substring(string, start, length?) → stringA slice of a string, indexed from 1 — not from 0, which is the mistake that shifts every result by one character.
- substring-before()substring-before(string, marker) → stringEverything before the first occurrence of a marker — and the empty string when the marker is absent, which is not an error you will notice.
- substring-after()substring-after(string, marker) → stringEverything after the first occurrence of a marker, and the empty string when the marker is absent. There is no split(): XPath 1.0 has no sequences.
- string-length()string-length(string?) → numberThe number of characters in a string, or of the context node with no argument. Characters, not bytes — a multi-byte character counts once.
- normalize-space()normalize-space(string?) → stringTrims leading and trailing whitespace and collapses internal runs to single spaces — the fix for text that looks equal and is not.
- translate()translate(string, from, to) → stringCharacter-by-character replacement, and the only way to fold case in XPath 1.0. A shorter third argument deletes the surplus rather than leaving it.
- concat()concat(string, string, string*) → stringJoins two or more strings, and requires at least two. There is no + for strings in XPath: 'a' + 'b' is arithmetic on two non-numbers, so it is NaN.
- string()string(object?) → stringConverts anything to a string — and for a node-set, takes only the first node in document order. Of an empty node-set it is '', not an error.
- sum()sum(node-set) → numberAdds the numeric values of every node in a set. One non-numeric node makes the whole result NaN, with no indication which node caused it.
- number()number(object?) → numberConverts to a number, producing NaN rather than an error when it cannot — which is how bad data passes silently.
- round(), floor() and ceiling()round(number) → numberThe three rounding functions. round() goes to the nearest integer, breaking ties upward — including for negatives.
- boolean() and not()boolean(object) → booleanConverts to true or false. A node-set is true when it is non-empty, which is why [not(x)] means 'has no x' — and why boolean('false') is true.
Axes · 9
- child:: axischild::name — or just nameImmediate children: the default axis, which is why it is almost never written out. It does not reach grandchildren, and never selects attributes.
- descendant:: and descendant-or-self::descendant::name — // is descendant-or-self::node()/Everything below a node, at any depth. // is shorthand for descendant-or-self, which is why //a//b can be expensive.
- parent:: axisparent::name — or ..The node directly above. Abbreviated .., and the way to select a node by what it contains. An attribute has a parent even though it is not a child.
- ancestor:: and ancestor-or-self::ancestor::nameEvery node above the context, up to the root. It is a reverse axis, so [1] means the nearest ancestor rather than the outermost one.
- attribute:: axisattribute::name — or @nameAn element's attributes, which live on their own axis and are never selected by a child step. Namespace declarations are not among them.
- following-sibling:: and preceding-sibling::following-sibling::nameSiblings after (or before) the context node. preceding-sibling:: is a reverse axis, which changes what [1] means.
- self:: axisself::name — or .The context node itself, abbreviated as a dot. Useful as a test: self::line is true only when the context node is a line, which filters a mixed node-set.
- following:: axisfollowing::nameEvery node after the context node's closing tag, across the rest of the document rather than only among its siblings.
- preceding:: axispreceding::nameEvery node before the context node's opening tag, excluding its ancestors. A reverse axis, so [1] means nearest.
Operators · 10
- node() testaxis::node()Matches every node kind on an axis: elements, text, comments, processing instructions and the root node. Indentation is text, so it returns more than you see.
- text() node testchild::text() — usually text()Selects text-node children, indentation whitespace included. Direct children only — descendant text needs .//text() or the element's string-value.
- comment() node testchild::comment() — usually comment()Selects XML comment nodes. Their delimiters are syntax; the node's string-value is only the content between them.
- processing-instruction() node testprocessing-instruction() or processing-instruction('target')Selects processing instructions, optionally restricted to one target such as xml-stylesheet. The target must be a string literal, not a QName.
- Predicates [ ]step[expression]Filters a node-set. Order matters, and a numeric predicate is not the same as a positional one applied afterwards.
- Union |node-set | node-setCombines two node-sets into one, in document order and without duplicates. Both operands must be node-sets, and the order you write them cannot matter.
- Comparison = and !=object = objectComparison against a node-set is existential, which makes != mean something other than 'not ='. Use not(a = b) when you actually mean negation.
- Arithmetic + - * div modnumber div numberDivision is div, not / — because / is the path separator. There is no string concatenation either: 'a' + 'b' is arithmetic, and therefore NaN.
- Path steps / and ///step/step — //stepA single slash is one level, a double slash any depth, and a leading slash makes the path absolute. Inside a predicate // is still absolute: use .// instead.
- Logical and / orboolean and boolean'and' binds tighter than 'or', and both are words rather than symbols — && and || do not parse. Parenthesise anything whose grouping is not obvious.
Get started
Bring order to the XML your team can't afford to ignore.
Create a free account and get a private workspace to search, validate, diff, and monitor your XML feeds, sitemaps, schemas, and vendor integrations.