Opened 4 months ago

Closed 3 months ago

#23487 closed enhancement (fixed)

Python Module Dependency Updates - certifi-2026.6.17 chardet-7.4.3 charset-normalizer-3.4.7 commonmark-0.9.2 editables-0.6 hatchling-1.30.1 idna-3.18 meson_python-0.20.0 msgpack-1.2.1 pathspec-1.1.1 pytz-2026.2 setuptools_rust-1.12.1 setuptools_scm-10.1.2 snowballstemmer-3.1.1 uv_build-0.11.24

Reported by: Douglas R. Reno Owned by: SecurityAdvisory
Priority: high Milestone: 98-Security
Component: BOOK Version: git
Severity: medium Keywords:
Cc:

Description

New versions of several Python module dependencies.

This contains two security updates - one to idna, and one to msgpack. The msgpack vulnerability is rated as High.

Change History (21)

comment:1 by Douglas R. Reno, 4 months ago

Owner: changed from blfs-book to Douglas R. Reno
Status: new → assigned

comment:2 by Douglas R. Reno, 4 months ago

certifi

A git log comparing the previous releases can be found here: ​https://github.com/certifi/python-certifi/compare/2026.02.25...2026.06.17

comment:3 by Douglas R. Reno, 4 months ago

chardet

6.0.0

6.0.0

Features

Unified single-byte charset detection: Instead of only having trained language 
models for a handful of languages (Bulgarian, Greek, Hebrew, Hungarian, Russian, Thai, 
Turkish) and relying on special-case Latin1Prober and MacRomanProber heuristics for 
Western encodings, chardet now treats all single-byte charsets the same way: every 
encoding gets proper language-specific bigram models trained on CulturaX corpus data. 
This means chardet can now accurately detect both the encoding and the language for all 
supported single-byte encodings.

38 new languages: Arabic, Belarusian, Breton, Croatian, Czech, Danish, Dutch, 
English, Esperanto, Estonian, Farsi, Finnish, French, German, Icelandic, Indonesian, 
Irish, Italian, Kazakh, Latvian, Lithuanian, Macedonian, Malay, Maltese, Norwegian, 
Polish, Portuguese, Romanian, Scottish Gaelic, Serbian, Slovak, Slovene, Spanish, 
Swedish, Tajik, Ukrainian, Vietnamese, and Welsh. Existing models for Bulgarian, Greek, 
Hebrew, Hungarian, Russian, Thai, and Turkish were also retrained with the new pipeline.

EncodingEra filtering: New encoding_era parameter to detect allows filtering by an 
EncodingEra flag enum (MODERN_WEB, LEGACY_ISO, LEGACY_MAC, LEGACY_REGIONAL, DOS, 
MAINFRAME, ALL) allows callers to restrict detection to encodings from a specific era. 
detect() and detect_all() default to MODERN_WEB. The new MODERN_WEB default should 
drastically improve accuracy for users who are not working with legacy data. The tiers 
are:

MODERN_WEB: UTF-8/16/32, Windows-125x, CP874, CJK multi-byte (widely used on the web)

LEGACY_ISO: ISO-8859-x, KOI8-R/U (legacy but well-known standards)

LEGACY_MAC: Mac-specific encodings (MacRoman, MacCyrillic, etc.)

LEGACY_REGIONAL: Uncommon regional/national encodings (KOI8-T, KZ1048, CP1006, etc.)

DOS: DOS/OEM code pages (CP437, CP850, CP866, etc.)

MAINFRAME: EBCDIC variants (CP037, CP500, etc.)

--encoding-era CLI flag: The chardetect CLI now accepts -e/--encoding-era to control 
which encoding eras are considered during detection.

max_bytes and chunk_size parameters: detect(), detect_all(), and UniversalDetector now 
accept max_bytes (default 200KB) and chunk_size (default 64KB) parameters for 
controlling how much data is examined. (#314, @bysiber)

Encoding era preference tie-breaking: When multiple encodings have very close confidence 
scores, the detector now prefers more modern/Unicode encodings over legacy ones.

Charset metadata registry: New chardet.metadata.charsets module provides structured 
metadata about all supported encodings, including their era classification and language 
filter.

should_rename_legacy now defaults intelligently: When set to None (the new default), 
legacy renaming is automatically enabled when encoding_era is MODERN_WEB.

Direct GB18030 support: Replaced the redundant GB2312 prober with a proper GB18030 
prober.

EBCDIC detection: Added CP037 and CP500 EBCDIC model registrations for mainframe 
encoding detection.

Binary file detection: Added basic binary file detection to abort analysis earlier on 
non-text files.

Python 3.12, 3.13, and 3.14 support (#283, @hugovk; #311)

GitHub Codespace support (#312, @oxygen-dioxide)

Fixes

Fix CP949 state machine: Corrected the state machine for Korean CP949 encoding 
detection. (#268, @nenw)

Fix SJIS distribution analysis: Fixed SJISDistributionAnalysis discarding valid second-
byte range >= 0x80. (#315, @bysiber)

Fix UTF-16/32 detection for non-ASCII-heavy text: Improved detection of UTF-16/32 
encoded CJK and other non-ASCII text by adding a MIN_RATIO threshold alongside the 
existing EXPECTED_RATIO.

Fix get_charset crash: Resolved a crash when looking up unknown charset names.

Fix GB18030 char_len_table: Corrected the character length table for GB18030 multi-byte 
sequences.

Fix UTF-8 state machine: Updated to be more spec-compliant.

Fix detect_all() returning inactive probers: Results from probers that determined 
"definitely not this encoding" are now excluded.

Fix early cutoff bug: Resolved an issue where detection could terminate prematurely.

Default UTF-8 fallback: If UTF-8 has not been ruled out and nothing else is above the 
minimum threshold, UTF-8 is now returned as the default.

Breaking changes

Dropped Python 3.7, 3.8, and 3.9 support: Now requires Python 3.10+. (#283, @hugovk)

Removed Latin1Prober and MacRomanProber: These special-case probers have been 
replaced by the unified model-based approach described above. Latin-1, MacRoman, and all 
other single-byte encodings are now detected by SingleByteCharSetProber with trained 
language models, giving better accuracy and language identification.

Removed EUC-TW support: EUC-TW encoding detection has been removed as it is extremely 
rare in practice.

LanguageFilter.NONE removed: Use specific language filters or LanguageFilter.ALL 
instead.

Enum types changed: InputState, ProbingState, MachineState, SequenceLikelihood, and 
CharacterCategory are now IntEnum (previously plain classes or Enum). LanguageFilter values changed from hardcoded hex to auto().

detect() default behavior change: detect() now defaults to 
encoding_era=EncodingEra.MODERN_WEB and should_rename_legacy=None (auto-enabled for 
MODERN_WEB), whereas previously it defaulted to considering all encodings with no legacy 
renaming.

Misc changes

Switched from Poetry/setuptools to uv + hatchling: Build system modernized with hatch-
vcs for version management.

License text updated: Updated LGPLv2.1 license text and FSF notices to use URL instead of mailing address. (#304, #307, @musicinmybrain)

CulturaX-based model training: The create_language_model.py training script was 
rewritten to use the CulturaX multilingual corpus instead of Wikipedia, producing higher 
quality bigram frequency models.

Language class converted to frozen dataclass: The language metadata class now uses 
@dataclass(frozen=True) with num_training_docs and num_training_chars fields replacing 
wiki_start_pages.

Test infrastructure: Added pytest-timeout and pytest-xdist for faster parallel test 
execution. Reorganized test data directories.

6.0.0.post1

6.0.0.post1

Fixed version number in chardet/version.py still being set to 6.0.0dev0. Otherwise 
identical to 6.0.0.

7.0.0

Ground-up, MIT-licensed rewrite of chardet. Same package name, same public API — drop-in 
replacement for chardet 5.x/6.x. Just way faster and more accurate!

Highlights:

MIT license (previous versions were LGPL)

96.8% accuracy on 2,179 test files (+2.3pp vs chardet 6.0.0, +7.7pp vs charset-
normalizer)

41x faster than chardet 6.0.0 with mypyc (28x pure Python), 7.5x faster than charset-
normalizer

Language detection for every result (90.5% accuracy across 49 languages)

99 encodings across six eras (MODERN_WEB, LEGACY_ISO, LEGACY_MAC, LEGACY_REGIONAL, DOS, MAINFRAME)

12-stage detection pipeline — BOM, UTF-16/32 patterns, escape sequences, binary 
detection, markup charset, ASCII, UTF-8 validation, byte validity, CJK gating, 
structural probing, statistical scoring, post-processing

Bigram frequency models trained on CulturaX multilingual corpus data for all supported 
language/encoding pairs

Optional mypyc compilation — 1.49x additional speedup on CPython

Thread-safe detect() and detect_all() with no measurable overhead; scales on free-threaded Python 3.13t+

Negligible import memory (96 B)

Zero runtime dependencies

Breaking changes vs 6.0.0:

detect() and detect_all() now default to encoding_era=EncodingEra.ALL (6.0.0 defaulted 
to MODERN_WEB)

Internal architecture is completely different (probers replaced by pipeline stages). 
Only the public API is preserved.

LanguageFilter is accepted but ignored (deprecation warning emitted)

chunk_size is accepted but ignored (deprecation warning emitted)

7.0.1

7.0.1
Fixes

Fixed false UTF-7 detection of SHA-1 git hashes (#324, fixing #323) — requirements files 
with VCS pins (e.g., +4bafdea3...) were misdetected as UTF-7, breaking tools like tox

Fixed _SINGLE_LANG_MAP missing aliases for single-language encoding lookup (e.g., big5 → 
big5hkscs)

Fixed PyPy TypeError in UTF-7 codec handling

Improvements

Retrained bigram models — 24 previously failing test cases now pass

Updated language equivalences for mutual intelligibility (Slovak/Czech, East Slavic + 
Bulgarian, Malay/Indonesian, Scandinavian languages)

7.1.0

Features

Added PEP 263 encoding declaration detection — # -*- coding: ... -*- and # 
coding=... declarations on lines 1–2 of Python source files are now recognized with 
confidence 0.95 (#249)

Added chardet.universaldetector backward-compatibility stub so that from 
chardet.universaldetector import UniversalDetector works with a deprecation warning 
(#341)

Fixes

Fixed false UTF-7 detection of ASCII text containing ++ or +word patterns (#332)

Fixed 0.5s startup cost on first detect() call — model norms are now computed during 
loading instead of lazily iterating 21M entries (#333)

Fixed undocumented encoding name changes between chardet 5.x and 7.0 — detect() now 
returns chardet 5.x-compatible names by default (#338)

Improved ISO-2022-JP family detection — recognizes ESC sequences for ISO-2022-
JP-2004 (JIS X 0213) and ISO-2022-JP-EXT (JIS X 0201 Kana)

Fixed silent truncation of corrupt model data (iter_unpack yielded fewer tuples instead 
of raising)

Fixed incorrect date in LICENSE

Performance

5.5x faster first-detect time (~0.42s → ~0.075s) by computing model norms as a side-
product of load_models()

~40% faster model parsing via struct.iter_unpack for bulk entry extraction (eliminates 
~305K individual unpack calls)

New API parameters

Added compat_names parameter (default True) to detect(), detect_all(), and 
UniversalDetector — set to False to get raw Python codec names instead of chardet 
5.x/6.x compatible display names

Added prefer_superset parameter (default False) — remaps legacy ISO/subset encodings to 
their modern Windows/CP superset equivalents (e.g., ASCII → Windows-1252, ISO-8859-1 → 
Windows-1252). This will default to True in the next major version (8.0).

Deprecated should_rename_legacy in favor of prefer_superset — a deprecation warning is 
emitted when used

Improvements

Switched internal canonical encoding names to Python codec names (e.g., "utf-8" instead 
of "UTF-8"), with compat_names controlling the public output format

Added lookup_encoding() to registry for case-insensitive resolution of arbitrary 
encoding name input to canonical names

Achieved 100% line coverage across all source modules (+31 tests)

Updated benchmark numbers: 98.2% encoding accuracy, 95.2% language accuracy on 2,510 
test files

Pinned test-data cloning to chardet release version tags for reproducible builds

7.2.0

chardet 7.2.0

Features

Added include_encodings and exclude_encodings parameters to detect(), detect_all(), 
and UniversalDetector — restrict or exclude specific encodings from the candidate set, 
with corresponding -i/--include-encodings and -x/--exclude-encodings CLI flags (#343)

Added no_match_encoding (default "cp1252") and empty_input_encoding (default 
"utf-8") parameters — control which encoding is returned when no candidate survives the 
pipeline or the input is empty, with corresponding CLI flags (#343)

Added -l/--language flag to chardetect CLI — shows the detected language (ISO 639-1 
code and English name) alongside the encoding (#342)

Fixes

Fixed null-separated ASCII data being misdetected as UTF-16-BE (#346, #347)

7.3.0

License

0BSD license — the project license has been changed from MIT to 0BSD, a maximally 
permissive license with no attribution requirement. All prior 7.x releases should also 
be considered 0BSD licensed as of this release.

Features

Added mime_type field to detection results — identifies file types for both binary (via 
magic number matching) and text content. Returned in all detect(), detect_all(), and 
UniversalDetector results. (#350)

New pipeline/magic.py module detects 40+ binary file formats including images, 
audio/video, archives, documents, executables, and fonts. ZIP-based formats (XLSX, DOCX, 
JAR, APK, EPUB, wheel, OpenDocument) are distinguished by entry filenames. (#350)

Bug Fixes

Fixed incorrect equivalence between UTF-16-LE and UTF-16-BE in accuracy testing — 
these are distinct encodings with different byte order, not interchangeable

Performance

Added 4 new modules to mypyc compilation (orchestrator, confusion, magic, ascii), 
bringing the total to 11 compiled modules

Capped statistical scoring at 16 KB — bigram models converge quickly, so large files 
no longer score the full 200 KB. Worst-case detection time dropped from 62ms to 26ms 
with no accuracy loss.

Replaced dataclasses.replace() with direct DetectionResult construction on hot 
paths, eliminating ~354k function calls per full test suite run

Build

Added riscv64 to the mypyc wheel build matrix — prebuilt wheels are now published 
for RISC-V Linux alongside existing architectures (#348, thanks @gounthar)

7.4.0

chardet 7.4.0

chardet 7.4.0 brings accuracy up to 99.3% (from 98.6% in 7.3.0) and significantly faster 
cold start thanks to a new dense model format.

What's New

Performance:

New dense zlib-compressed model format (v2) drops cold start (import + first detect) 
from ~75ms to ~13ms with mypyc

Accuracy (98.6% → 99.3%):

Eliminated train/test data overlap via content fingerprinting

Added MADLAD-400 and Wikipedia as supplemental training sources

Improved non-ASCII bigram scoring: high-byte bigrams are now preserved during training 
and weighted by per-bigram IDF

Encoding-aware substitution filtering (substitutions only apply for characters the 
target encoding can't represent)

Increased training samples from 15K to 25K per language/encoding pair

Bug fixes:

Added dedicated structural analyzers for CP932, CP949, and Big5-HKSCS (these were 
previously sharing their base encoding's byte-range analyzer, missing extended ranges)

7.4.1

7.4.1

Bug Fixes

BOM-prefixed UTF-16/32 input now returns utf-16/utf-32 instead of utf-16-le/utf-16-
be/utf-32-le/utf-32-be. The endian-specific codecs don't strip the BOM on decode, so 
callers were getting a stray U+FEFF at the start of their text. BOM-less detection is 
unchanged. (#364, #365)

7.4.2

7.4.2

Patch release: fixes a crash on short inputs and closes a bunch of WHATWG/IANA alias gaps.

Bug Fixes

Fixed RuntimeError: pipeline must always return at least one result on ~2% of all 
possible two-byte inputs (e.g. b"\xf9\x92"). Multi-byte encodings like CP932 and Johab 
could score above the structural confidence threshold on very short inputs, but then 
statistical scoring would return nothing, leaving an empty result list instead of 
falling through to the fallback. (#367, #368, thanks @jasonwbarnett)

Improvements

Added ~90 encoding aliases from the WHATWG Encoding Standard and IANA Character Sets 
registry so that <meta charset> labels like x-cp1252, x-sjis, dos-874, csUTF8, and the 
cswindows* family all resolve correctly through the markup detection stage. Every alias 
was driven by a failing spec-compliance test, not speculative. (#366)

Added a spec-compliance test suite covering Python decode round-trips for all 86 
registry encodings, WHATWG label resolution, IANA preferred MIME names, and Unicode/RFC 
conformance (BOM sniffing, UTF-8 boundary cases, UTF-16 surrogate pairs). This is the 
test suite that would have caught the 7.4.1 BOM bug before release. (#366)

7.4.3

Patch release: fixes a crash when input contains null bytes inside a <meta charset> 
declaration.

Bug Fixes

Fixed ValueError: embedded null character crash when input contained a <meta 
charset> declaration with a null byte in the encoding name (e.g. b'<meta 
charset="\x00utf-8">'). codecs.lookup() raises ValueError on embedded nulls, and 
lookup_encoding() was only catching LookupError. Also added defensive ValueError catches 
in _validate_bytes() and _to_utf8() for completeness. (#369, thanks @DRMacIver for the 
report)

comment:4 by Douglas R. Reno, 4 months ago

charset-normalizer

3.4.5

3.4.5 (2026-03-06)
Changed

Update setuptools constraint to setuptools>=68,<=82.

Raised upper bound of mypyc for the optional pre-built extension to v1.19.1

Fixed

Add explicit link to lib math in our optimized build. (#692)

Logger level not restored correctly for empty byte sequences. (#701)

TypeError when passing bytearray to from_bytes. (#703)

Misc

Applied safe micro-optimizations in both our noise detector and language detector.

Rewrote the query_yes_no function (inside CLI) to avoid using ambiguous licensed code.

Added cd.py submodule into mypyc optional compilation to reduce further the performance impact.

3.4.6

3.4.6 (2026-03-15)

Changed

Flattened the logic in charset_normalizer.md for higher performance. Removed 
eligible(..) and feed(...) in favor of feed_info(...).

Raised upper bound for mypy[c] to 1.20, for our optimized version.

Updated UNICODE_RANGES_COMBINED using Unicode blocks v17.

Fixed

Edge case where noise difference between two candidates can be almost insignificant. 
(#672)

CLI --normalize writing to wrong path when passing multiple files in. (#702)

Misc

Freethreaded pre-built wheels now shipped in PyPI starting with 3.14t. (#616)

3.4.7

3.4.7 (2026-04-02)
Changed

Pre-built optimized version using mypy[c] v1.20.

Relax setuptools constraint to setuptools>=68,<82.1.

Fixed

Correctly remove SIG remnant in utf-7 decoded string. (#718) (#716)

comment:5 by Douglas R. Reno, 4 months ago

commonmark

Deprecate package. Use markdown-it-py instead.

Remove testing on python 3.4.

comment:6 by Douglas R. Reno, 4 months ago

editables

Release 0.6

Add a new "self_replace" strategy for map (and name the old strategy "import_hook"). 
Based on an idea by Daniel Tang in #40.

Rename the generated .pth file to _editable_impl_<project>.pth and document that it is 
possible to customise the file names used.

Rework the documentataion, replacing the "use cases" section with an expanded and less 
opinionated "scope" section.

Test suite improvements.

comment:7 by Douglas R. Reno, 4 months ago

hatchling

1.29.0

Hatchling v1.29.0

Fixed:

    Source Date Epoch no longer fails when set to date before 1980.

1.30.1

Hatchling v1.30.1

Fixed

    Default core metadata version kept at 2.4 until more tools support 2.5

comment:8 by Douglas R. Reno, 4 months ago

idna

3.12

3.12 (2026-04-21)

Update to Unicode 17.0.0.

Issue a deprecation warning for the transitional argument.

Added lazy-loading to provide some performance improvements.

Removed vestiges of code related to Python 2 support, including segmentation of data 
structures specific to Jython.

3.13

3.13 (2026-04-22)

    Correct classification error for codepoint U+A7F1

3.14

3.14 (2026-05-10)

Removed opportunity to process long inputs into quadratic time by rejecting oversize 
inputs up-front. Closes a bypass of the CVE-2024-3651 mitigation. [CVE-2026-45409]

3.15

3.15 (2026-05-12)

Enforce DNS-length cap on individual labels early in check_label, short-circuiting 
contextual-rule processing for oversized input while staying compatible with UTS 46 
usage.

Tidy core helpers: hoist bidi category sets to module-level frozensets (avoiding 
per-codepoint list construction), simplify length checks, and reuse the shared 
_unicode_dots_re from idna.core in the codec module.

Use raise ... from err for proper exception chaining and switch internal string 
formatting to f-strings.

Allow flit_core 4.x in the build backend.

Expand the ruff lint set (flake8-bugbear, flake8-simplify, pyupgrade, perflint) and 
apply the surfaced fixes; pin lint CI to Python 3.14.

Add Dependabot configuration for GitHub Actions.

Convert README and HISTORY from reStructuredText to Markdown.

Reference CVE-2026-45409 for the 3.14 advisory in place of the initial GHSA identifier.

3.16

3.16 (2026-05-22)

Add a command-line interface (python -m idna, also available as the idna script). 

Encodes or decodes one or more domains supplied as arguments or on standard input, with 
options to select A-label or U-label output and control error handling.

Raise the minimum supported Python version to 3.9

Various code quality improvements

3.17

3.17 (2026-05-28)

Substantial 75% reduction in memory usage through new data structures and some 
optimization in processing speed.

Added a general 1024-character input length cap to the public validation, 
conversion, and codec entry points. This is well above any legitimate domain or label 
and guards against pathological inputs.

3.18

3.18 (2026-06-02)

When decoding a domain, add a display argument that will pass through invalid labels 
rather than raising an exception.

comment:9 by Douglas R. Reno, 4 months ago

meson_python

0.20.0

0.20.0

    Add support for targeting the Android platform.

comment:10 by Douglas R. Reno, 4 months ago

Summary: Python Module Dependency Updates - certifi-2026.6.17 chardet-7.4.3 charset-normalizer-3.4.7 commonmark-0.9.2 editables-0.6 hatchling-1.30.1 idna-3.18 meson_python-0.12.0 msgpack-1.2.1 pathspec-1.1.1 pytz-2026.2 setuptools_rust-1.12.1 setuptools_scm-10.1.2 snowballstemmer-3.1.1 uv_build-0.11.24 → Python Module Dependency Updates - certifi-2026.6.17 chardet-7.4.3 charset-normalizer-3.4.7 commonmark-0.9.2 editables-0.6 hatchling-1.30.1 idna-3.18 meson_python-0.20.0 msgpack-1.2.1 pathspec-1.1.1 pytz-2026.2 setuptools_rust-1.12.1 setuptools_scm-10.1.2 snowballstemmer-3.1.1 uv_build-0.11.24

Correct the meson_python version number.

comment:11 by Douglas R. Reno, 4 months ago

msgpack

1.2.0

1.2.0

Release Date: 2026-06-11

Support free threaded Python. #654, #686

Dropped support for Python 3.9. #656

Fix missing error checks in C code. #665, #666, #667, #672

Fix strict_map_key option didn't work for object_pairs_hook. #673

Increase DEFAULT_RECURSE_LIMIT of Unpacker to 1024. #676

Fix memory leak when Unpacker returns error for invalid input. #671

Fix Packer.pack_ext_type() ignored autoreset option. #663

Fix Timestamp.from_datetime() returning wrong value for pre-epoch datetimes. #662

Fix use-after-free in unpackb() and Unpacker.unpack() for non-contiguous input. #677

Fix possible memory leak when calling Unpacker.__init__() several times. #687

1.2.1

Fix a segfault when calling Unpacker.unpack() or Unpacker.skip() after an unpacking 
failure. But note that reusing the same Unpacker instance after an unpacking failure is 
not supported. Please create a new Unpacker instance instead. GHSA-6v7p-g79w-8964

comment:12 by Douglas R. Reno, 4 months ago

pathspec

1.1.1 (2026-04-26)

Improvements:

    Improved type checking with mypy and pyright.

Bug fixes:

    Fixed typing on PathSpec[TPattern] to PathSpec[TPattern_co].

    Added missing variant type-hint type[Pattern] to PathSpec.from_lines() parameter 
pattern_factory.

    Fixed possible type error when using + and += operators on PathSpec.

1.1.0 (2026-04-22)

New features:

    Issue #108: Specialize pattern type for PathSpec as PathSpec[TPattern] for better 
debugging of PathSpec().patterns.

Bug fixes:

    Issue #93: Git discards invalid range notation. GitIgnoreSpecPattern now discards 
patterns with invalid range notation like Git.

    Pull #106: Fix escape() not escaping backslash characters.

Improvements:

    Pull #110: Nicer debug print outs (and str for regex pattern).

comment:13 by Douglas R. Reno, 4 months ago

pytz

A git log comparison between 2025.2 and 2026.2 can be found here: ​https://github.com/stub42/pytz/compare/release_2025.2...release_2026.2

comment:14 by Douglas R. Reno, 4 months ago

setuptools_rust

1.12.1 (2026-03-26)

    Migrate to trusted publishing. #581
    Strip target suffix for cargo-zigbuild compatibility. #534

comment:15 by Douglas R. Reno, 4 months ago

setuptools_scm

10.0.0

Removed

Drop Python 3.8 and 3.9 support. Minimum Python version is now 3.10. (#1228)

Added

setuptools-scm now depends on vcs-versioning for core version inference logic. This 
enables other build backends to use the same version inference without setuptools 
dependency. (#1228)

Version files (write_to and version_file) are now written to the build directory
during build_py instead of the source tree during version inference.
This enables installing packages from read-only source directories (e.g., Bazel builds).

Path transformation is automatically applied for src/ layouts - a configured path like
src/mypackage/_version.py is correctly written to mypackage/_version.py in the
build directory based on the package_dir configuration.

To restore the old behavior of writing version files at inference time (useful for
development workflows), set the environment variable SETUPTOOLS_SCM_WRITE_TO_SOURCE=1. (#1252)

Fixed

Fix issue #1231: Don't warn about tool.setuptools.dynamic.version conflict when only 
using file finder without version inference. (#1231)

Miscellaneous

Refactored should_infer from method to standalone function for better code organization. 
(#1228)

Updated mypy version template test to use uvx, ensuring generated version files remain 
compatible with Python 3.8+ consumers. (#1228)

Refactored TestBuildPackageWithExtra into parametrized function with custom INI-based 
decorator for cleaner test data specification. (#1228)

Internal refactoring: modernized type annotations, improved CLI type safety, and 
enhanced release automation infrastructure. (#1228)

10.0.1

Miscellaneous

Simplify release tag creation to use a single createRelease API call instead of 
separate createTag/createRef/createRelease calls, avoiding dangling tag objects on 
partial failures. (#release-pipeline)

10.0.2

Fixed

Fix version file not generated for editable installs. Version files are now written 
to the source tree by default during inference (restoring pre-10.x behavior), and also 
registered as build_py outputs so strict editable installs include them in the 
persistent auxiliary directory. Set SETUPTOOLS_SCM_WRITE_TO_SOURCE=0 to disable source-
tree writing (e.g., for read-only source directories). (#1298)

10.0.3

Fixed

Remove monorepo-only ../vcs-versioning/src from build-system.backend-path so sdists 
install under PEP 517 (paths must stay inside the source tree). (#1306)

Miscellaneous

Add griffecli to test dependencies so the API stability check keeps working after the 
Griffe CLI was split into a separate package. (#1310)

10.0.4

Fixed

Anchor get_version in setup.py with relative_to and fallback_root so SCM fallbacks (e.g. 
PKG-INFO) do not resolve against the wrong directory when the build cwd is the workspace 
or repo root. (#1302)

Enter GlobalOverrides for SETUPTOOLS_SCM when using setuptools_scm.get_version / 
_get_version, avoiding implicit context warnings for direct API callers. (#1314)

Miscellaneous

Upgrade pre-commit hooks (Ruff, mypy, codespell), align locked Ruff with hooks, and add 
Ruff per-file configuration for setuptools_scm re-export modules. (#1311)

10.0.5

Fixed

Allow dump_version() deprecation warning to be silenced by passing scm_version=None. 
(#1286)

Remove [tool.uv.sources] from setuptools-scm/pyproject.toml to fix sdist builds outside 
the workspace — the workspace root already declares the source mapping for development. 
(#1330)

10.1.0

Added

Add backward-compatible shims in setuptools_scm.git, setuptools_scm.hg, 
setuptools_scm.hg_git, and setuptools_scm.scm_workdir so that external code calling 
get_scm_version(config) or run_describe(config) with an explicit Configuration continues 
to work. The shim automatically wires _config and VcsEnvironment onto the workdir. 
(#compat-shims)

Write scm_version.json and scm_file_list.json into egg-info directories during 
egg_info, enabling sdist fallback version inference when no VCS is present. Add 
ScmEggInfoMixin for workdir-based file finding in find_sources(). (#egg-info-metadata)

Add write_to_source pyproject.toml option to control whether version files are 
written to the source tree. When unset, a deprecation warning advises setting it 
explicitly before the default changes in a future major release. The 
SETUPTOOLS_SCM_WRITE_TO_SOURCE environment variable overrides this setting. (#1301)

Adopt the workdir-centric pipeline from vcs-versioning: version discovery now 
follows an explicit env → config → workdir → version chain instead of relying on ambient 
globals and parse entry points. The egg_info command writes scm_version.json and 
scm_file_list.json metadata so sdists can infer versions without a VCS checkout. 
Requires vcs-versioning >= 2.0.0.dev0. (#1378)

Fixed

Fix worktree file listing test to expect relative paths from the file finder. The test 
now passes on Linux; Windows remains xfail due to a subprocess limitation with worktree 
directories. (#620)

Remove the _warn_on_old_setuptools() check that incorrectly warned when a custom 
build-backend caused setuptools.__version__ to return the project version instead of 
setuptools' version. The minimum setuptools version is now enforced via build-system 
requirements. (#1192)

Wrap version in setuptools.sic() when normalize = false to prevent setuptools from 
re-normalizing the version after our hook returns. This preserves CalVer zero-padding 
(e.g. 2024.01.05) and other non-canonical version strings in dist.metadata.version. 
(#1354)

Skip writing non-package version files to build_lib, fixing incorrect inclusion of root-
level version files in wheels. (#1364)

Documentation

Rewrite the GitHub Actions CI/CD example to use a dedicated build job
(via build-and-inspect-python-package) and OIDC Trusted Publishers
instead of building in publishing jobs with long-lived API tokens. (#1215)

10.1.1

Fixed

Update CI to use PyPy 3.11 as cryptography has no PyPy 3.10 build available (#1421)

10.1.2

Fixed

Fix DeprecationWarning leak by threading VcsEnvironment through 
VersionInferenceConfig and using env.make_reader() in _should_write_to_source. (#1424)

comment:16 by Douglas R. Reno, 4 months ago

snowballstemmer

3.1.0

Python
------

* Bug fixes:

  + Fix `algorithms()` when forwarding to PyStemmer.  It looks like this has
    never worked as the code has been like this since it was merged, and we
    were forwarding to a method which PyStemmer doesn't provide and never seems
    to have provided.

  + stemwords.py: Make -i and -o optional.  The command syntax already
    suggested they were, but actually we gave an error if they were omitted.

  + Fix code generated for string-$ (which isn't used by any of the algorithms
    we currently ship).

  + Fix `->` to work when the slice is empty - previously it incorrectly
    signalled `f` for this case.  Luckily this case is not exercised by any
    current algorithms (#242)

  + Remove deprecated licence classifier which now triggers a deprecation
    warning from Python's setuptools.  We already specify the licensing in the
    now preferred way via `license=` with a SPDX licence expression.

* Optimisations:

  + Optimise single-character string literal checks in the same way we already
    do for C.  This seems to be measurably faster (tested with Turkish which
    has lots of single character literal tests).

  + Groupings are now implemented via a Python set, or a string for small
    groupings.

  + Eliminate use of exception in code generated for `or`.  We can instead wrap
    the code in a loop and use `break`.

  + Eliminate use of exception in `goto` and `gopast`.  We can just use `break`
    here to exit the `while` loop we're also inside and move the `except` from
    the previous `try` onto the `while`.

  + Avoid using a temporary for `hop` with a constant argument as benchmarking
    with timeit shows this is faster.

  + Optimise string test by using startswith()/endswith() with suitable
    start/end parameters which avoids creating a temporary substring and avoids
    an explicit limit check.  This speeds up artificial testcases consisting of
    `goto 'the'` by 10%.

  + Optimise among when all actions are `<-` with a literal string.  We now
    generate a single call to slice_from() with the argument obtained by
    indexing into an array of literal strings.  See #227.

  + Reduce overhead of code to forward to PyStemmer, both when forwarding and
    when using the pure Python stemmers.

  + Reuse exception classes much more.  This reduces the number of labN classes
    we need by 142 over all the current stemmers.

  + Change slice_check() to assert its conditions.  In C we must not perform
    string slicing if slice_check() fails because that could result in writing
    outside of the allocated buffer, but it's not problematic in this way for
    Python, and the situations which slice_check() checks for should only
    happen with a Snowball program containing logic errors, or for bugs in the
    Snowball compiler or its runtime (or possibly in the Python interpreter,
    OS, hardware, etc).  Therefore assert() seems an appropriate choice.

* Code quality:

  + Use _ as dummy loop variable.  We don't use the loop variable's value, and
    the loop itself tracks the current iteration so generating nested loops
    using `_` as the loop variable works correctly.

  + Avoid mysterious gaps in the numbering of variables in the generated code.
    This was already done for the other languages, but I missed Python it
    seems.

  + Avoid generating unused lab0 class for a Snowball program which doesn't use
    any failure labels.

  + Avoid generating a blank line at start of the body of a Snowball `loop`.

  + stemwords.py: Replace deprecated `codecs.open()` with built-in `open()`.
    Patch from Dmitry Shachnev.

* Documentation:

  + Remove unnecessary semicolons from Python code in docs.

* Other changes:

  + Remove Python 2 support.  We stopped officially supporting it in Snowball
    2.1.0, but now we've actually stripped out support.  Versions of Python ≥
    3.3 continue to be supported.  Patch from Dmitry Shachnev (#212).

3.1.1

Python
------

* Other changes:

  + Skip classifier for Sesotho which isn't yet in the official list of
    trove classifiers.  Patch from Dmitry Shachnev (#289).

  + Add classifier to indicate support for Python 3.14.

in reply to:  15 comment:17 by Douglas R. Reno, 4 months ago

Replying to Douglas R. Reno:

setuptools_scm

10.0.0

Removed

Drop Python 3.8 and 3.9 support. Minimum Python version is now 3.10. (#1228)

Added

setuptools-scm now depends on vcs-versioning for core version inference logic. This 
enables other build backends to use the same version inference without setuptools 
dependency. (#1228)

Version files (write_to and version_file) are now written to the build directory
during build_py instead of the source tree during version inference.
This enables installing packages from read-only source directories (e.g., Bazel builds).

Path transformation is automatically applied for src/ layouts - a configured path like
src/mypackage/_version.py is correctly written to mypackage/_version.py in the
build directory based on the package_dir configuration.

To restore the old behavior of writing version files at inference time (useful for
development workflows), set the environment variable SETUPTOOLS_SCM_WRITE_TO_SOURCE=1. (#1252)

Fixed

Fix issue #1231: Don't warn about tool.setuptools.dynamic.version conflict when only 
using file finder without version inference. (#1231)

Miscellaneous

Refactored should_infer from method to standalone function for better code organization. 
(#1228)

Updated mypy version template test to use uvx, ensuring generated version files remain 
compatible with Python 3.8+ consumers. (#1228)

Refactored TestBuildPackageWithExtra into parametrized function with custom INI-based 
decorator for cleaner test data specification. (#1228)

Internal refactoring: modernized type annotations, improved CLI type safety, and 
enhanced release automation infrastructure. (#1228)

10.0.1

Miscellaneous

Simplify release tag creation to use a single createRelease API call instead of 
separate createTag/createRef/createRelease calls, avoiding dangling tag objects on 
partial failures. (#release-pipeline)

10.0.2

Fixed

Fix version file not generated for editable installs. Version files are now written 
to the source tree by default during inference (restoring pre-10.x behavior), and also 
registered as build_py outputs so strict editable installs include them in the 
persistent auxiliary directory. Set SETUPTOOLS_SCM_WRITE_TO_SOURCE=0 to disable source-
tree writing (e.g., for read-only source directories). (#1298)

10.0.3

Fixed

Remove monorepo-only ../vcs-versioning/src from build-system.backend-path so sdists 
install under PEP 517 (paths must stay inside the source tree). (#1306)

Miscellaneous

Add griffecli to test dependencies so the API stability check keeps working after the 
Griffe CLI was split into a separate package. (#1310)

10.0.4

Fixed

Anchor get_version in setup.py with relative_to and fallback_root so SCM fallbacks (e.g. 
PKG-INFO) do not resolve against the wrong directory when the build cwd is the workspace 
or repo root. (#1302)

Enter GlobalOverrides for SETUPTOOLS_SCM when using setuptools_scm.get_version / 
_get_version, avoiding implicit context warnings for direct API callers. (#1314)

Miscellaneous

Upgrade pre-commit hooks (Ruff, mypy, codespell), align locked Ruff with hooks, and add 
Ruff per-file configuration for setuptools_scm re-export modules. (#1311)

10.0.5

Fixed

Allow dump_version() deprecation warning to be silenced by passing scm_version=None. 
(#1286)

Remove [tool.uv.sources] from setuptools-scm/pyproject.toml to fix sdist builds outside 
the workspace — the workspace root already declares the source mapping for development. 
(#1330)

10.1.0

Added

Add backward-compatible shims in setuptools_scm.git, setuptools_scm.hg, 
setuptools_scm.hg_git, and setuptools_scm.scm_workdir so that external code calling 
get_scm_version(config) or run_describe(config) with an explicit Configuration continues 
to work. The shim automatically wires _config and VcsEnvironment onto the workdir. 
(#compat-shims)

Write scm_version.json and scm_file_list.json into egg-info directories during 
egg_info, enabling sdist fallback version inference when no VCS is present. Add 
ScmEggInfoMixin for workdir-based file finding in find_sources(). (#egg-info-metadata)

Add write_to_source pyproject.toml option to control whether version files are 
written to the source tree. When unset, a deprecation warning advises setting it 
explicitly before the default changes in a future major release. The 
SETUPTOOLS_SCM_WRITE_TO_SOURCE environment variable overrides this setting. (#1301)

Adopt the workdir-centric pipeline from vcs-versioning: version discovery now 
follows an explicit env → config → workdir → version chain instead of relying on ambient 
globals and parse entry points. The egg_info command writes scm_version.json and 
scm_file_list.json metadata so sdists can infer versions without a VCS checkout. 
Requires vcs-versioning >= 2.0.0.dev0. (#1378)

Fixed

Fix worktree file listing test to expect relative paths from the file finder. The test 
now passes on Linux; Windows remains xfail due to a subprocess limitation with worktree 
directories. (#620)

Remove the _warn_on_old_setuptools() check that incorrectly warned when a custom 
build-backend caused setuptools.__version__ to return the project version instead of 
setuptools' version. The minimum setuptools version is now enforced via build-system 
requirements. (#1192)

Wrap version in setuptools.sic() when normalize = false to prevent setuptools from 
re-normalizing the version after our hook returns. This preserves CalVer zero-padding 
(e.g. 2024.01.05) and other non-canonical version strings in dist.metadata.version. 
(#1354)

Skip writing non-package version files to build_lib, fixing incorrect inclusion of root-
level version files in wheels. (#1364)

Documentation

Rewrite the GitHub Actions CI/CD example to use a dedicated build job
(via build-and-inspect-python-package) and OIDC Trusted Publishers
instead of building in publishing jobs with long-lived API tokens. (#1215)

10.1.1

Fixed

Update CI to use PyPy 3.11 as cryptography has no PyPy 3.10 build available (#1421)

10.1.2

Fixed

Fix DeprecationWarning leak by threading VcsEnvironment through 
VersionInferenceConfig and using env.make_reader() in _should_write_to_source. (#1424)

Note that I needed to add the vcs_versioning module for this. I used version 2.1.2 which was most current as of today, and that module has no extra dependencies.

comment:18 by Douglas R. Reno, 4 months ago

uv_build is a slimmed down version of the uv package, but shares the same respository. The release notes are massive and can be found at ​https://github.com/astral-sh/uv/blob/0.11.24/CHANGELOG.md - but I suspect 99% of this isn't applicable to us since we're just using the build backend.

comment:19 by Douglas R. Reno, 4 months ago

Owner: changed from Douglas R. Reno to SecurityAdvisory
Status: assigned → new

Fixed at 9d7d1afe9ed7718b9218dc157e6ff482e1aa0ce4

Added the XML file for the vcs_versioning module at 8b4f3a4f678acb384fd76d046d75802cc454a8d3

Reassigning to SecurityAdvisory for an advisory for idna/msgpack to be filed.

comment:20 by Bruce Dubbs, 3 months ago

Milestone: 13.1 → 98-Security

comment:21 by Bruce Dubbs, 3 months ago

Resolution: → fixed
Status: new → closed

Added new advisory sa-13.0-143.

Note: See TracTickets for help on using tickets.