Skip to content

Issue #169 probe: docs-only PR must receive the required CI Status check - #174

Merged
jeremymanning merged 201 commits into
masterfrom
docs/169-probe
Aug 23, 2026
Merged

Issue #169 probe: docs-only PR must receive the required CI Status check#174
jeremymanning merged 201 commits into
masterfrom
docs/169-probe

Conversation

@jeremymanning

Copy link
Copy Markdown
Member

Deliberately touches only README.md. If the required CI Status and Tests Status contexts report on this pull request, then #169's fix (the status-check job no longer path-filtered away) is proven on the wire and #169 can close. Merge or close this PR either way — its content is one HTML comment.

jeremymanning and others added 30 commits August 19, 2026 13:58
Locates the conflicting original design element (#64: "No local file
support - cannot access local modules, data files, or custom code",
listed as a defect with a "File Packaging" solution direction), and
records the contrary position in #10 where syncing data/results was
explicitly rejected -- so the feature is asymmetric: inputs up, outputs
only on request.

Establishes from the code what exists: SFTP transport is wired
(executor_connections.py:146/156) but only for the payload; file-reference
detection (dependency_analysis.py:249) and packaging (file_packaging.py:319)
are fully orphaned; the cluster_* API is read-only; shared-filesystem
detection exists at filesystem.py:113.

Proposes 5 independently shippable phases, each verified against real
SLURM/SSH hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…tore failure

`clustrix/enhanced_notebook_widget.py` (139 statements, 0.00% coverage) was
unreachable. `%%remote` goes through `notebook_magic_core.display_config_widget`
-> `modern_notebook_widget.display_modern_widget`; nothing in `clustrix/`
imported the enhanced widget, it is not in `clustrix/__init__.py`, and
`docs/source/api/notebook_magic.rst` never mentioned it. Its only importers
were two real-world visual tests, both removed here:

- `test_enhanced_widget_html_output` rendered the deleted widget and nothing
  else.
- `test_widget_comparison_report` existed to diff it against the modern
  widget; with one side gone its only remaining behaviour -- dumping the
  modern widget's HTML -- duplicates `test_modern_widget_html_output`.

`AuthenticationManager`, `validate_cluster_auth` and `validate_ssh_key_auth`
are kept: `modern_notebook_widget.py` still calls all three.

`SecureCredentialManager.store_credential` returned `False` after logging a
warning, which is indistinguishable from a credential that was written and
then lost. It now raises `NotImplementedError` naming the item and the
supported alternative (`~/.clustrix/.env` + `clustrix.credential_manager`).
It had no callers, so nothing breaks. `ensure_secure_environment()` is
deleted outright -- zero callers anywhere in the repository.

`notebook_magic_widget.py` is deliberately untouched: it is public API
(`clustrix.notebook_magic.__all__`) documented as a compatibility shim, and
removing it is an API break that needs its own decision.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
docs/source/notebooks/local_parallel_comparison.ipynb benchmarks, with
every number produced by executing the notebook:

- a plain call vs @cluster(cores=N) under cluster_type="local" (1.00x --
  the local branch of the decorator calls the function in this
  interpreter; cores has no scheduler to reach)
- CPU-bound work through LocalExecutor: processes ~5.6x, threads ~1.0x
- I/O-bound work through LocalExecutor: threads and processes both ~6x
- fan-out too small to pay for a pool (speedup below 1.0x)
- what auto_parallel does to a qualifying function locally

It also states which pool is used and what selects it: use_threads on
LocalExecutor, filled in by choose_executor_type when left as None --
picklability of the function and every argument first, then a substring
scan of the function's source for I/O names, then processes.

Registered in the Interactive Notebooks toctree.
The content of this change landed in bb7fbd0 ("Add measured
local-execution tutorial notebook") because a concurrent commit on this
branch picked up files that were staged at the time. The code is correct
and present; only the commit message is wrong. This empty commit records
what bb7fbd0 actually contains for #148, since rewriting a branch other
agents are committing to is not safe.

What changed:

Every test that opened a paramiko connection called
set_missing_host_key_policy(paramiko.AutoAddPolicy()) -- 37 sites across
29 files -- which trusts whatever host key a server offers on first
contact. clustrix/ssh_security.py was written to close exactly that hole
in production code; the test suite kept its own copy of it, so every SSH
connection the real-world suite made was still unverified.

All 37 sites now call configure_host_key_policy(client), passing the
surrounding ClusterConfig at the 9 sites where one is in scope.

This is deliberately NOT behaviour-preserving. The default policy is
"reject", so a test pointed at a host absent from the runner's
known_hosts now raises HostKeyVerificationError instead of connecting.
That is the point.

NOT VERIFIED BY EXECUTION: no affected test was run. They all require
real SSH credentials and live cluster or localhost sshd access. The
change is verified only by static means -- import/collection of all 156
real-world tests, black, flake8, and the unit suite.

Also:
- .github/workflows/real-world-tests.yml gains a "Populate known_hosts"
  step (ssh-keyscan) before the test steps, covering localhost/127.0.0.1
  -- the sshd that job installs, and the only host the CI run reaches --
  plus any CLUSTRIX_TEST_{SSH,SLURM}_HOST[_2] that are set. No repository
  secret supplies an external cluster hostname; the step does not pretend
  one exists.
- docs/REAL_WORLD_TESTING.md and docs/testing_guidelines.md document the
  known_hosts prerequisite with the exact ssh-keyscan command.
- tests/unit/test_no_autoadd_policy.py fails if AutoAddPolicy reappears in
  any Python source outside clustrix/ssh_security.py and
  tests/unit/test_host_key_policy.py. It runs in ordinary CI, asserts the
  scan can still see the two allowed uses (so it cannot pass vacuously),
  and exercises its own failure path against a planted violation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
The suite's worst mock tests are the three named in #117. All three
asserted values the test itself had supplied to a stub:

  * test_executor.py test_execute_command set
    mock_stdout.read.return_value = b"command output" and asserted
    stdout == "command output" -- it tested Python assignment.
  * test_ssh_automation.py test_validate_ssh_key_success mocked
    paramiko.SSHClient, called validate_ssh_key("/path/to/key") and
    asserted True. No key, no host, no authentication.
  * test_filesystem.py test_remote_ls fed a MagicMock
    b"file1.txt\nfile2.py\nsubdir/\n" and asserted those names came
    back. It tested str.split and rstrip.

Replaced with real execution. tests/ssh_server.py is paramiko's server
side in process: a real socket on 127.0.0.1, real host keys, real
password and public-key authentication, exec channels run by a real
shell, and a real SFTP subsystem over a real directory. clustrix runs
against it unmodified -- no production code knows anything about it.

Also converted, in the same three files: connect/disconnect, SLURM
status and cancellation, function-data serialization, SSH key
generation, deployment, detection and lookup, setup_ssh_keys end to
end, and the filesystem exists/stat/connection-failure paths.

Two real defects surfaced and are now pinned by assertions rather than
mocked past:

  * setup_ssh_keys can never report connection_tested for a key it
    generated: step 5 uses detect_existing_ssh_key, which only tries
    the six standard key names, and the generated key is never one.
  * ClusterFilesystem._remote_stat issues GNU-only `stat -c`, so a BSD
    remote host raises FileNotFoundError for a file that exists.

Mock occurrences measured with the issue's own regex: 596 across 33 of
177 files before (not the 2,513 the issue claims), 525 across 31 after,
with the remaining matches in these three files being prose in
docstrings only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…tations

- Drop every reference to what older versions of clustrix did.
- Restate the data-mover boundary as current behaviour and point at #151.
- Add a limitation for @cluster(cores=N) not parallelizing locally (#152).
- Correct the joblib comparison and the unknown-keyword warning behaviour.
- file_packaging and dependency_analysis: say plainly that nothing on the
  execution path calls them, instead of claiming they replace pickling.
- filesystem: state that the cluster_* utilities are read-only.
- decorator: cores does nothing on the local path; drop before/after notes.
- config, notebook_magic, ssh_setup, configuration: present tense only.
…buckets

Implements the design in @jeremymanning's comment on #151, which supersedes the
issue body's SFTP-into-_stage plan: a DataPackage object built outside the
cluster call, carrying a local path and a private HF bucket path, or the bytes
themselves when the data is small or force_local=True. Passed to a @cluster
function as an ordinary argument and dereferenced on the worker on demand.

Departures from the issue body, all recorded in notes/data-package-implementation.md:
new clustrix/staging.py rather than cluster_put/cluster_get in filesystem.py;
one uuid-keyed folder per package rather than a content-addressed refcounted
store; no reaper, no TTL, no automatic cleanup of any kind -- deletion is always
an explicit act by the user, per the owner's follow-up.

Kept from the issue body: declaration never inference, the three size bands,
digest verification, atomic .partial+rename, path confinement with escapes
rejected not clamped, refusal of credential-shaped paths, 0600/0700 modes.

Digests are computed locally and travel inside the upload-only function payload,
so they never pass through the store -- which is why staged data needs no HMAC.

Streaming deferred to #155. file_packaging.py is superseded, not adopted: it
packages source-inferred dependencies, the inference this design rejects.

Verified against real things: real files, real @cluster execution, the real
paramiko SSH server over real SFTP through the shipped ConnectionManager, a real
private HuggingFace repo (upload, download, verify, delete, list -- kilobytes
only), and a real subprocess for the pickle-survives-a-fresh-interpreter path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
- @cluster(cores=N) does not parallelize locally (#152); say so up front.
- Add a Data section: nothing but the pickled call travels, #151 is the plan.
- Correct the mock-usage count (20 of 152, not 42 of 215) and the local
  backend's description (calling process, not local processes).
- Drop the complexity-analysis claim: no such analysis exists.
- Remove every 'was removed in v0.2.0' framing outside the backend table.
Also removes complexity analysis and function flattening from the source-
availability note (both were deleted), corrects the mock-module count, and
drops the 'this file previously said' framing.
Corrects 'Local - multiprocessing and threading' (it is neither, #152),
drops the non-existent performance/resource monitoring claim, and replaces
the emoji-heavy SSH summary with accurate guidance on work directories and
per-call SSH connections.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…real known_hosts

Two defects found by running the tests repeatedly rather than once.

The upload was one HuggingFace commit per file. The Hub rate-limits commits to
256 per hour per account, so a package of any size could exhaust a user's quota
on its own -- and repeated test runs did exactly that. It is now a single
create_commit carrying every file plus the manifest, which is cheaper and also
atomic: either the whole package lands or none of it does, so the partial-upload
rollback path is gone rather than merely untested.

The real-HuggingFace tests are marked real_world so the standard
-m 'not real_world' command does not spend the owner's quota on every run. They
are still never mocked; without a token they skip and the remote half is simply
unverified.

Separately, connecting to the in-process SSH server with auto_add made paramiko
rewrite the developer's entire ~/.ssh/known_hosts non-atomically, which corrupted
it in 7 of 15 runs and then broke every subsequent SSH connection. Reported as
#157; the fixture here redirects HOME so these tests cannot touch the real file.
15 consecutive runs now pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
The 'planned rather than present' wording became false when 9719280 landed
clustrix/staging.py. Verified data_package() builds and round-trips a real
file before changing the claim. The cluster_* helpers stay read-only and
file_packaging.py is still off the execution path.
Where a package was last unpacked is true of one machine at one moment. Carrying
it into the pickle lets a worker -- or the same machine a week later -- find a
stale directory at that path and reuse it. Everything else on the object is
durable; that one field is not, so __getstate__ drops it.

Also removes a no-op try/finally left in _atomic_write, and threads the caller's
config through _discard_local_cache so delete() clears the cache directory the
package actually used rather than whatever the global config points at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…e dead with-block reference

test_no_autoadd_policy.py scans for the string, not the call, so explaining the
defect by name in a comment failed the scan. Reworded; the scan is right to be
blunt. Also removes a mention of a context-manager form that no longer exists
and says why there deliberately is not one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
A 429 from the Hub surfaced as a bare HfHubHTTPError out of huggingface_hub,
which tells the caller nothing about what clustrix was doing or what to do next.
It now says which repo, which package, that the limit is per hour per account,
and -- the part that actually matters -- that nothing was uploaded, because a
package is a single commit so a refused commit leaves no partial state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
check_docs_examples' `--include-notes` widened the scan to every prose file
under docs/, including session logs and design analyses -- records of what
someone believed at the time -- and it *ran* their code blocks. One makes a
live AWS API call. Executing a year-old example to find out whether it
still parses is not a trade worth making.

Notes are now inventoried rather than executed: a `never_execute` flag on
TargetFile, carried across the subprocess boundary, routing every block
through the static verifier.

    default:          143 checked, 110 executed,  33 static
    --include-notes:  243 checked, 110 executed, 133 static

The executed count is unchanged -- the extra 100 blocks are read and
syntax-checked, never run -- and the flag still does its job, reporting 24
broken examples in those historical files.

Also drops MIGRATION.md from the target list; that file has been deleted.
It documented a repository reorganization ("Before: Cluttered Repository /
After: Clean Organization"), which is version-to-version prose of the kind
being removed everywhere else, and it was wrong besides -- line 70 asserted
that `from clustrix import ClusterConfig` does not work, when it does and
the name is in `__all__`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
Nine remote commands interpolated caller-supplied paths and glob patterns
with no shlex.quote anywhere in the module, and every one is reachable from
exported public API. Six sites move to SFTP, where a path is a path: ls,
stat, exists, isdir, isfile and glob. du is walked over SFTP too, because
its `du -sb` was GNU-only in exactly the way `stat -c` was. find keeps a
shell -- walking a tree over SFTP would cost a round trip per directory --
with both its directory and its pattern shlex.quote'd; quoting is what stops
the *shell* expanding the pattern, find still expands it itself. count_files
is now find, counted, rather than a second command.

_remote_stat no longer runs GNU-only `stat -c` behind `2>/dev/null`, which
reported FileNotFoundError for files that existed on any BSD or macOS host.
A failing remote command now logs its exit status and stderr instead of
returning empty output the caller reads as an absence.

find_ssh_keys matched six exact filenames, none of which a clustrix-generated
key can ever have, so setup_ssh_keys could not verify the key it had just
deployed and a second run did not notice one existed. Discovery now also
matches id_<type>_clustrix[_<suffix>], and nothing else in ~/.ssh.

tests/unit/test_filesystem_injection.py proves against the real in-process
SSH server that ; && | $() `` ' newline and a leading dash are treated as
filenames, that the pre-fix command strings really did execute them, that
globbing still globs, and that no new unquoted interpolation can appear in
the module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…et inline-only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
Measured before this change on a developer machine: **1,191 of the 1,223
entries** in the real known_hosts were loopback junk left by test runs. 32
were genuine. Two runs at once interleave their appends and corrupt the
file, after which unrelated tests fail with InvalidHostKey -- and so does
the developer's own ssh.

Two independent causes, both fixed.

**1. Only some test files isolated $HOME.** Several had grown their own
`isolated_home` fixture; everything else wrote wherever the developer
lives. `tests/conftest.py` now has an autouse `isolate_home` covering every
test that exists and every test anyone writes later, which is the only
version of this that stays true. `USERPROFILE` is set alongside `HOME`
because that is what `Path.home()` reads on Windows.

**2. OpenSSH ignores $HOME.** It resolves "~" from the passwd database, so
`ssh-copy-id` appended to the real file no matter what the environment
said. Confirmed directly:

    $ HOME=/tmp/xxxx ssh -G -o StrictHostKeyChecking=accept-new host
    userknownhostsfile /Users/jmanning/.ssh/known_hosts

That is not only a test problem. clustrix's Python half computes
`~/.ssh/known_hosts` from `$HOME` -- `add_host_key`, and
`ssh_security._load_known_hosts` -- while its subprocess half used the
passwd home. Anywhere the two differ (a container, `sudo -u`, a login node
with a relocated home) clustrix verifies host keys against a file it is not
writing to. `_user_known_hosts_path()` is now the single definition, and
`ssh-copy-id` is passed `-o UserKnownHostsFile=` explicitly so both halves
agree.

Verified: the full non-billable suite leaves known_hosts byte-identical.

    1345 passed, 18 skipped, 0 failed

One assertion rewritten, and the old one was wrong rather than merely
stale: `test_deploy_public_key_ssh_copy_id_success` pinned the exact
ssh-copy-id argv, which is the right thing to assert -- it now also asserts
the known_hosts option is present and that the path follows `$HOME`, since
without that the isolation is nominal.

The 1,191 stale entries already in the developer's file are theirs to
remove; nothing here touches them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
tests/conftest.py now gives every test a throwaway $HOME, which is right -- the
suite had left 1,191 junk loopback entries in the developer's real known_hosts.
It also put the token written by `hf auth login` out of reach, so all eight
real-HuggingFace tests failed for want of a credential.

The token is now resolved at import time, before any fixture runs, and handed to
the hf_config fixture via monkeypatch. That reaches exactly the tests that need
it -- and the subprocess one of them spawns, which is how a worker gets a token
in production anyway -- rather than putting a credential in the environment of
every test in the suite. On a machine with no credentials they skip, and the
skip reason says plainly that the remote half went unverified rather than
passed.

Drops this file's own isolated-known-hosts fixture, now redundant with conftest.

Also documents two things a user should learn before they happen rather than
after: clustrix creates a private dataset repo in their account from a derived
name, and deleting packages never deletes that repo -- only folders inside it,
because hf_data_repo may point at a repo they own and care about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
Records what landed, the settings changes, the three open filesystem defects
the red team found, the documentation errors not yet fixed, and where to
resume.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
- local_cache_dir is read (staging.py materialises data packages under it);
  moved out of the "no effect" table into the paths table.
- The environment-variable section claimed three variables are read. Replaced
  the count with the categories: config location, import behaviour,
  credentials, and the internal channel clustrix uses to talk to its own
  remote code. The point it was making -- no CLUSTRIX_<FIELD> overlay sets a
  config field -- is unchanged and still true.
- Documented job_wait_timeout, hf_data_repo, stage_inline_max_bytes,
  stage_warn_bytes and stage_max_bytes, so the page's completeness claim
  holds against dataclasses.fields(ClusterConfig).
- ssh_port is read by validate_cluster_auth as well as auth_manager.
- An unsupported cluster_type raises at configuration time, not submit time.
- default_queue and @cluster(queue=...) are resolved into the job config and
  never read; listed as having no effect.

The removed-backend error message no longer frames itself as a change between
versions. tests/test_config.py asserts on the new wording and additionally
asserts no version string appears in it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
executor_schedulers.py's module docstring said PBS and SGE "were removed in
v0.2.0", which is the version-to-version framing being taken out everywhere
else: a reader wants to know what clustrix does, not what it stopped doing.

It now states the current position -- no PBS/Torque or SGE submission ships,
because neither has been verified against a real scheduler of that kind --
and keeps the issue numbers.

`grep -rn "v0\.2\.0" clustrix/` is now empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
Commit 95b523e introduced `_user_known_hosts_path()` in ssh_utils and said
it was "the single definition". It was not: ssh_security.py:92 went on
computing `Path(os.path.expanduser("~/.ssh/known_hosts"))` independently,
so the module that *verifies* host keys and the module that *writes* them
each had their own copy of the rule.

Two copies of a rule like this drift, and the drift is invisible until the
two disagree -- which is exactly the failure the original commit was
fixing, one level up.

The definition now lives in ssh_security.py, the lower-level module, and
ssh_utils imports it. Import direction was already ssh_utils -> ssh_security,
so no cycle. Verified they are the same function object:

    from clustrix.ssh_security import user_known_hosts_path
    from clustrix.ssh_utils import _user_known_hosts_path
    same object: True

    $ grep -rn 'expanduser("~/.ssh/known_hosts")' clustrix/ | wc -l
    0

Found by an adversarial review of 95b523e.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…atch local

Three defects an adversarial review of 411b6aa found. The injection fix
itself holds and is untouched.

1. The anti-regression guard only inspected assignments to `cmd*`-named
   variables and arguments to `exec_command`, so `self._run_remote(f"ls -1
   {full_path}")` -- the module's own primary helper, fully exploitable --
   reported zero violations, along with seven other spellings. It now
   tracks taint: shell sinks are discovered by following a function's own
   parameters into `exec_command` to a fixpoint (which is how `_run_remote`
   is found rather than named), and a value is clean only if it is a
   literal, a `shlex.quote` result, or built from clean values. Rebinding
   the name `shlex` is a violation in its own right.

2. `glob("*/")` no longer returned directories only, and `_local_glob` --
   the oracle -- did. `_remote_glob` now runs `glob.glob`'s own algorithm
   over SFTP instead of a hand-rolled component split, which also fixes
   absolute patterns, `glob("")`, `glob(".")` and unnormalised `..`.

3. `_remote_du` counted a symlink as its target and re-descended through a
   link to an ancestor forever (320 bytes / 32 phantom files on a tree
   holding 10 bytes / 1 file). It now classifies entries the way
   `os.walk(followlinks=False)` plus `os.path.getsize` do, which is also
   why it terminates: no symlink is followed, so no directory is reached
   twice.

Also: remote `permissions` was malformed below three octal digits --
`oct(mode & 0o777)[-3:]` gives "0o0", "0o7", "o77" where local gives "000",
"007", "077".

tests/test_filesystem.py gains a TestLocalAndRemoteAgree class that asserts
the two implementations answer identically over the same real directory, so
the asymmetry cannot drift back. Fifteen of its assertions fail against the
pre-fix module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012gTBDPK16HUZ3kHQ2QyjuU
…kernel without it

The CI rerun got past nbformat and died at the next gate: 'Kernel died
before replying to kernel_info'. nbclient launches kernels through
ipykernel, which the dev extra never declared; local environments had it
transitively and so passed.
…onda skip for the invariants

The ubuntu-3.11 CI job failed every HOME-find test while 3.10/3.12 and
macos passed identical code, so the failure carries its inputs now: PATH,
HOME, fixture existence and mode, rc, stdout, stderr.

The submission_invariants guard (nothing but the fixture's conda may be
reachable from the test account) skips when the runner image itself ships
/usr/bin/conda -- what a shared runner installs is not something a commit
can fix, and clustrix's own search order is pinned hermetically by the
emitted-script tests.
…nostics

Three CI findings from the probe rerun:

- docs-test: three docstrings put literal spans inside bold spans (the
  new public-api page pulled them into rendered HTML). Bold closed and
  reopened around each literal; markup checker back to OK across 43
  pages.
- windows legs died at exactly timeout-minutes: nbclient kernel launches
  hang there. _execute_notebook raises NotebookPlatformUnavailable on
  win32; the checker reports those notebooks held-back with the reason,
  mirroring cluster-required. The same notebooks execute on linux and
  macos.
- the named-env diagnostic attached its env via a mock-only attribute;
  it now rides on the result itself.

2786 passed / 0 failed; sphinx -W clean; markup OK; examples 239/239.
File exists, mode right, PATH stubbed -- and [ -f ] still says no on
ubuntu. Speculation is exhausted; bash -x reports which comparison
failed and what _clustrix_base held at that moment.
…on via timeout

bash -x from the CI diagnostics caught it exactly: after sourcing the
site's conda.sh -- which defines conda as a shell FUNCTION --
_clustrix_conda_works ran its check as 'timeout 10 conda --version'.
timeout is an external binary that execs files: it never sees the
function, and instead ran whichever conda FILE came first on PATH. On
GitHub's ubuntu runners that file is broken, so the job died claiming no
conda existed while one sat sourced in the very shell asking; on real
profile.d-initialised clusters the probe has been returning empty all
along, silently disabling conda discovery.

The helpers now ask the shell directly when  says
function (no wrapper -- functions cannot be exec'd), and keep the
timeout bound only for foreign executables. The pass-through timeout in
the named-env fixture reproduces the runner shape on any machine (RED
first: 9 failures with the stub+timeout combo), and the 19 byte-for-byte
goldens are regenerated from the fixed generator.

2787 passed / 0 failed.
…nnot read as empty

The ubuntu CI leg caught a second platform face of the unmeasurable-
job.err fix: GNU wc on Linux answers '0' for a directory and exits 0,
so a job.err that was actually a directory reported 'running' there
while BSD wc refused it into the unknown branch. Measuring through the
redirection (wc -l < path) fails the redirection on every platform for
a directory, landing in the same honest except; regular files are
unaffected.
… 0 for a directory through a stdin redirect

The redirect form was still not enough on Linux: opening a directory
read-only succeeds there, and GNU wc reports 0 lines for what it could
not read. test -f gates the measurement so an unmeasurable job.err
lands in the honest except on every platform; regular files count as
before.
The 5-minute step bound was set when tests/unit held ~350 fast tests;
the suite has tripled since and now executes notebooks and real
subprocesses, so the step was being cut at 60% regardless of what had
already passed. 15 minutes bounds today's suite with headroom.
Windows CI ran 271 failures across every suite that leans on
tests/ssh_server.py: exec requests run through a POSIX shell, permission
bits and key generation are Unix-shaped, and paramiko's server mode
needs socket semantics Windows denies. The server raises pytest.skip on
win32 with that reason, so dependent tests skip instead of failing in
bulk; bringing local SSH testing to Windows is its own piece of work.
…lence assertion listens to the right logger

- The win32 guard in tests/ssh_server.py fires during collection for
  suites that build the server at import time; without
  allow_module_level=True that is a collection error x17 instead of a
  clean skip.
- test_a_listening_socket asserted caplog.records == [] -- every logger.
  paramiko's server thread races a banner warning into that list on
  loaded runners (the prober hangs up mid-handshake, which is the probe
  working). The assertion now scopes to clustrix.notebook_magic_widget,
  which is the logger the test names.
The Windows leg collected cleanly after the ssh_server guard but 111
assertions still failed across twelve modules whose subject is POSIX
itself -- Unix permission bits, bash-run emitted scripts,
getcwd/permission edges, runner-shaped workflow contexts. Each module
now skips on win32 with its reason stated; unskipping selectively is
follow-up work per module.

2787 passed / 0 failed locally; black/flake8/mypy clean.
- test_config_loading_with_encoding_issues writes NUL and
  invalid-encoding bytes; Windows text mode mangles them before the YAML
  reader sees the error the test pins.
- test_comprehensive_edge_case_suite exercises POSIX permission, chmod
  and shell edges that do not exist on NTFS.
…s headroom

The macos legs died at timeout-minutes=30 with the suite only 15% done:
measured 14.4 tests/min vs ubuntu's 253 -- and 1,482 of the 1,780 seconds
spent sat in one file, TestLocalAndRemoteAgree. Every case is a real
SSH round-trip against the in-process server, and macOS CI runners pay
Gatekeeper verification on every process spawn.

That class now skips where it hurts (darwin AND CI); developer macOS and
two other platforms keep the full oracle coverage. The test job's bound
moves to 45 minutes so a slow leg reports pass/fail instead of being
executed for time.
@jeremymanning
jeremymanning merged commit 8d0f1c4 into master Aug 23, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant