#21779 closed enhancement (fixed)

lxml-6.0.0 (Python module)

Reported by: Bruce Dubbs Owned by: Bruce Dubbs
Priority: normal Milestone: 12.4
Component: BOOK Version: git
Severity: medium Keywords:
Cc:

Description

New major version.

Change History (3)

comment:1 by Bruce Dubbs, 15 months ago

Owner: changed from blfs-book to Bruce Dubbs
Status: new → assigned

comment:2 by Bruce Dubbs, 15 months ago

Summary: lxml-6.0.0 → lxml-6.0.0 (Python module)

lxml changelog - 6.0.0 (2025-06-26)

Features added

  • GH463: lxml.html.diff is faster and provides structurally better diffs.
  • GH405: The factories Element and ElementTree can now be used in type hints.
  • GH448: Parsing from memoryview and other buffers is supported to allow zero-copy parsing.
  • GH437: lxml.html.builder was missing several HTML5 tag names.
  • GH458: CDATA can now be written into the incremental xmlfile() writer.
  • A new parser option decompress=False was added that controls the automatic input decompression when using libxml2 2.15.0 or later. Disabling this option by default will effectively prevent decompression bombs when handling untrusted input. Code that depends on automatic decompression must enable this option. Note that libxml2 2.15.0 was not released yet, so this option currently has no effect but can already be used.
  • The set of compile time / runtime supported libxml2 feature names is available as etree.LIBXML_COMPILED_FEATURES and etree.LIBXML_FEATURES. This currently includes catalog, ftp, html, http, iconv, icu, lzma, regexp, schematron, xmlschema, xpath, zlib.

Bugs fixed

  • GH353: Predicates in .find*() could mishandle tag indices if a default namespace is provided.
  • GH272: The head and body properties of lxml.html elements failed if no such element was found. They now return None instead.
  • Tag names provided by code (API, not data) that are longer than INT_MAX could be truncated or mishandled in other ways.
  • .text_content() on lxml.html elements accidentally returned a "smart string" without additional information. It now returns a plain string.

Other changes

  • Support for Python < 3.8 was removed.
  • Parsing directly from zlib (or lzma) compressed data is now considered an optional feature in lxml. It may get removed from libxml2 at some point for security reasons (compression bombs) and is therefore no longer guaranteed to be available in lxml.

As of this release, zlib support is still normally available in the binary wheels but may get disabled or removed in later (x.y.0) releases. To test the availability, use "zlib" in etree.LIBXML_FEATURES.

  • The Schematron class is deprecated and will become non-functional in a future lxml version. The feature will soon be removed from libxml2 and stop being available.
  • GH438: Wheels include the arm7l target.
  • GH465: Windows wheels include the arm64 target.
  • Binary wheels use the library versions libxml2 2.14.4 and libxslt 1.1.43. Note that this disables direct HTTP and FTP support for parsing from URLs. Use Python URL request tools instead (which usually also support HTTPS). To test the availability, use "http" in etree.LIBXML_FEATURES.
  • Windows binary wheels use the library versions libxml2 2.11.9, libxslt 1.1.39 and libiconv 1.17. They are now based on VS-2022.
  • Built using Cython 3.1.2.
  • The debug methods MemDebug.dump() and MemDebug.show() were removed completely. libxml2 2.13.0 discarded this feature.

comment:3 by Bruce Dubbs, 15 months ago

Resolution: → fixed
Status: assigned → closed

Fixed at commits

00e44cf666 Update to nettle-3.10.2.
b5c802424f Update to lxml-6.0.0 (Python module).
40a4cf0123 Update to fontconfig-2.17.0.
Note: See TracTickets for help on using tickets.