Ffile2fix
Sign in Get started

How to fix XML parsing error: not well-formed (invalid token)

xml.parsers.expat.ExpatError: not well-formed (invalid token): line 14, column 32

The parser found a character or sequence that breaks XML's basic rules, so it stops at the line and column shown without reading further. The most common culprits are an unescaped & or <, bytes that are not valid in the declared encoding, and invisible control characters.

Also appears as: XML Parsing Error: not well-formed Location: https://example.com/feed.xml Line Number 14, Column 32: · simplexml_load_string(): Entity: line 14: parser error : xmlParseEntityRef: no name · This page contains the following errors: error on line 14 at column 32: EntityRef: expecting ';' · simplexml_load_file(): parser error : Input is not proper UTF-8, indicate encoding ! Bytes: 0xE9 0x20 0x6C 0x65

Common causes

  • A raw & in text or URLs (e.g. ?a=1&b=2 or 'Tom & Jerry') instead of &amp;
  • A literal < in text content, or HTML pasted into an element without CDATA
  • Text saved as Windows-1252/Latin-1 while the XML declares or defaults to UTF-8
  • Control characters (0x00-0x1F except tab, newline, carriage return) copied from Word or a database
  • HTML entities such as &nbsp; or &copy; that are not defined in XML
  • Unquoted attribute values or mismatched quotes in attributes

How to fix it

  1. Go to the reported position. Open the file in an editor that shows line and column and look at the exact character. For remote feeds, save the response to a file first (curl -o feed.xml URL).
  2. Escape special characters. Replace & with &amp; and < with &lt; in text and attribute values. In code, use an XML library or htmlspecialchars($text, ENT_XML1 | ENT_QUOTES, 'UTF-8') instead of string concatenation.
  3. Wrap markup in CDATA. If an element must contain HTML, wrap it in <![CDATA[ ... ]]> so the parser does not interpret it (the sequence ]]> must not appear inside).
  4. Fix the encoding. Save the file as UTF-8, or make the declaration match the real encoding, e.g. <?xml version="1.0" encoding="ISO-8859-1"?>. 'Input is not proper UTF-8' always means this.
  5. Remove control characters. Strip invalid characters before writing XML, e.g. preg_replace('/[^\x{9}\x{A}\x{D}\x{20}-\x{D7FF}\x{E000}-\x{FFFD}]/u', '', $text) in PHP.
  6. Replace HTML entities. Use numeric references (&#160; for a non-breaking space, &#169; for ©) or the real UTF-8 characters instead of HTML-only named entities.

PHP (build XML safely)

$doc = new DOMDocument('1.0', 'UTF-8');
$item = $doc->createElement('item');
$title = $doc->createElement('title');
$title->appendChild($doc->createTextNode('Tom & Jerry <Classic>')); // escaped automatically
$item->appendChild($title);
$doc->appendChild($item);
header('Content-Type: application/xml; charset=UTF-8');
echo $doc->saveXML();

How to stop it happening again

  • Generate XML with DOMDocument, XMLWriter or a library, never by concatenating strings
  • Keep data UTF-8 end to end and declare it in the XML prolog
  • Validate feeds and sitemaps after every change to their templates

Frequently asked questions

My WordPress feed or sitemap shows this error. What should I check?

Look at the exact position: often a post title or excerpt contains a stray control character or unescaped content from a plugin. Also check for output before <?xml, which causes a separate 'XML declaration allowed only at the start' error.

Is &nbsp; valid in XML?

No. XML only predefines &amp; &lt; &gt; &quot; and &apos;. Use &#160; or the actual character instead, unless a DTD defines nbsp.

Can I make the parser ignore the error?

XML parsers are required to stop on well-formedness errors. libxml (used by PHP's DOMDocument) has a recover mode, but it silently drops or changes data, so fix the source instead.