Read values from an XML file
Goal: Pull elements, text, and attributes out of an XML file such as a Maven pom.xml or an RSS feed, and print them as a Markdown table or list.
Prerequisites: The built-in xml module. Read the file with -I xml, which parses it into a tree of {tag, attributes, children, text} dicts. The example also uses the built-in csv module to render a table.
Query
The dependencies of a pom.xml as a Markdown table:
$ mq -I xml 'import "xml" | import "csv" | xml::xml_find_all("dependency") | map(fn(d): {"group": xml::xml_text(xml::xml_find(d, "groupId")), "artifact": xml::xml_text(xml::xml_find(d, "artifactId")), "version": xml::xml_text(xml::xml_find(d, "version"))};) | csv::csv_to_markdown_table()' pom.xml
Input (pom.xml)
<?xml version="1.0" encoding="UTF-8"?>
<project>
<modelVersion>4.0.0</modelVersion>
<artifactId>demo-app</artifactId>
<version>1.2.0</version>
<dependencies>
<dependency>
<groupId>org.slf4j</groupId>
<artifactId>slf4j-api</artifactId>
<version>2.0.9</version>
</dependency>
<dependency>
<groupId>com.google.guava</groupId>
<artifactId>guava</artifactId>
<version>33.0.0-jre</version>
<scope>test</scope>
</dependency>
</dependencies>
</project>
Output
| group | artifact | version |
| --- | --- | --- |
| org.slf4j | slf4j-api | 2.0.9 |
| com.google.guava | guava | 33.0.0-jre |
Filter by an attribute
xml_attr(element, name) returns an attribute’s value, or None when it is not set. This lists the id of every <user> whose role is dev:
$ echo '<users><user id="1" role="admin"/><user id="2" role="dev"/><user id="3" role="dev"/></users>' | mq -I xml 'import "xml" | xml::xml_find_all("user") | filter(fn(u): xml::xml_attr(u, "role") == "dev";) | map(fn(u): xml::xml_attr(u, "id");)'
["2", "3"]
Turn an RSS feed into a list
$ curl -s https://example.com/feed.xml | mq -I xml 'import "xml" | xml::xml_find_all("item") | map(fn(i): "- [" + xml::xml_text(xml::xml_find(i, "title")) + "](" + xml::xml_text(xml::xml_find(i, "link")) + ")";) | join("\n")'
With a feed that has two <item> entries:
- [v0.8.5](https://example.com/v0.8.5)
- [v0.8.4](https://example.com/v0.8.4)
Notes
xml_find_all(tree, tag)returns every element namedtagin document order, including the root itself.xml_findreturns only the first match, orNonewhen nothing matches. Because matching is at any depth,xml_find(tree, "version")on thepom.xmlabove returns the project’s own1.2.0, since it comes before the dependency versions.xml_text(element)joins the text of an element and all its descendants with single spaces. For<user><name>Alice</name><email>[email protected]</email></user>it returnsAlice [email protected].- Keep only some elements by testing a child:
filter(fn(d): xml::xml_text(xml::xml_find(d, "scope")) == "test";)keeps the test-scoped dependencies of thepom.xmlabove. Elements without a<scope>do not match. -I xmlis shorthand for reading with-I rawand callingxml::xml_parsefirst. Use the longer form when the XML is a string inside other data, such as a field of a JSON document.- For path-style queries such as
//user[@role="dev"]/name, see Query XML with XPath. To go from XML to JSON and back, see Convert between XML and JSON. - For RSS 2.0 and Atom feeds specifically, the feed.mq extension module parses both into one shape.