Improve KeyboardInterrupt handling during Arrow C stream reads in Python bindings by kosiew · Pull Request #1 · nuno-faria/datafusion-python

kosiew · 2025-11-29T12:42:33Z

Fix test_arrow_c_stream_interrupted in Upgrade to Datafusion 51 apache/datafusion-python#1311

Rationale for this change

After upgrading to the latest DataFusion version, we observed that long‑running Python operations backed by Arrow C streams were no longer reliably interruptible via KeyboardInterrupt (e.g. Ctrl+C).

The root cause is that simply calling py.check_signals() is not sufficient to trigger CPython's signal handling for pending signals; it only raises an exception if a signal has already been processed by a previous Python API call. When our async utility sits in a Rust/Tokio loop without executing any Python code, pending signals (like SIGINT) may never be processed, so KeyboardInterrupt is not delivered to user code.

This PR fixes the problem by:

Forcing the Python interpreter to run a trivial no‑op statement periodically from the async loop so that pending signals are processed.
Tightening and modernizing the Python test so it actually interrupts the thread performing the Arrow C stream read, rather than assuming the main thread is doing the read. This makes the test robust and aligned with the new behavior.

Together, these changes restore responsive KeyboardInterrupt behavior for long‑running DataFusion/Arrow C stream operations in the Python bindings after the DataFusion upgrade.

What changes are included in this PR?

Rust async signal handling (src/utils.rs)
- Import std::ffi::CString to safely construct C string literals for the Python C API.
- In future_into_py, replace the bare py.check_signals() call inside the tokio::select! loop with a small closure that:
  - Constructs a CString for the Python code "pass".
  - Calls py.run with this code, causing the interpreter to execute a no‑op Python statement.
  - Then calls py.check_signals() to raise KeyboardInterrupt (or other signal‑related exceptions) if a signal was processed while running the code.
- This ensures that pending signals are actually processed while the Rust future is waiting, making async operations interruptible again.
Python test for Arrow C stream interruption (python/tests/test_dataframe.py)
- Rework test_arrow_c_stream_interrupted to:
  - Introduce a read_started event to signal when the background thread begins consuming the RecordBatchReader.
  - Track the read_thread_id using threading.get_ident() inside the read thread, instead of assuming that the main thread is doing the read.
  - Use a shared read_exception list to capture the outcome of the read operation (e.g. KeyboardInterrupt, timeout, or unexpected exception).
- Add a read_stream helper function that:
  - Sets read_thread_id.
  - Sets the read_started event to inform the interrupt thread that the read is in progress.
  - Calls reader.read_all() and records the result:
    - If the read unexpectedly completes, it pushes a RuntimeError("Read completed without interruption") into read_exception.
    - If KeyboardInterrupt is raised, it records that.
    - Any other exception is also captured for assertion.
- Update trigger_interrupt to:
  - Wait for read_started with a max_wait_time timeout.
  - Use PyThreadState_SetAsyncExc to raise KeyboardInterrupt specifically in the read_thread_id instead of the main thread.
  - Reset the exception state and raise a RuntimeError if the injection fails.
- Start both the read thread and the interrupt thread, then:
  - join the read thread with a 10‑second timeout and fail the test if it is still alive (indicating a hang).
  - Assert that an exception was captured.
  - Assert that the captured exception is either a KeyboardInterrupt or clearly contains "KeyboardInterrupt" in its message, to handle cases where it might be wrapped.
  - Finally, join the interrupt thread with a short timeout.

Overall, the test now:

Targets the thread actually performing the blocking read.
Is resilient to timing issues.
Explicitly verifies that the interruption is manifested as a KeyboardInterrupt.

Are these changes tested?

Yes.

The existing test_arrow_c_stream_interrupted has been rewritten to more accurately model the real‑world scenario of interrupting a blocking Arrow C stream read running in a background thread.
The test validates that:
- The read operation does not hang indefinitely (it must complete or be interrupted within the timeout).
- A KeyboardInterrupt (or an exception clearly indicating KeyboardInterrupt) is raised when we inject the signal into the read thread.

No additional tests were required beyond this updated test, as it directly exercises the new behavior introduced by the future_into_py changes when integrated with the Python runtime.

Test in Jupyter Notebook

Are there any user-facing changes?

Yes, but they are behavioral improvements rather than API changes:

Python users running long‑running queries or scans via DataFusion/Arrow C streams should once again be able to interrupt those operations with Ctrl+C and receive a KeyboardInterrupt promptly.
There are no changes to public function signatures, modules, or configuration options; this is a bug fix in the underlying async and signal‑handling behavior.

No breaking changes to public APIs are introduced by this PR. The only observable change is more reliable and responsive handling of KeyboardInterrupt during long‑running operations.

…parate thread

Updated wait_for_future to surface pending Python exceptions by executing bytecode during signal checks, ensuring that asynchronous interrupts are processed promptly. Enhanced PartitionedDataFrameStreamReader to cancel remaining partition streams on projection errors or Python interrupts, allowing for clean iteration stops. Added regression tests to validate interrupted Arrow C stream reads and improve timing for RecordBatchReader.read_all cancellations.

…g and readability

…k for KeyboardInterrupt

…meStreamReader

This reverts commit 784929d.

kosiew · 2025-11-29T12:49:40Z

The Jupyter notebook is saved in commit 784929d for testing.

* Add AGENTS.md and enrich __init__.py module docstring Add python/datafusion/AGENTS.md as a comprehensive DataFrame API guide for AI agents and users. It ships with pip automatically (Maturin includes everything under python-source = "python"). Covers core abstractions, import conventions, data loading, all DataFrame operations, expression building, a SQL-to-DataFrame reference table, common pitfalls, idiomatic patterns, and a categorized function index. Enrich the __init__.py module docstring from 2 lines to a full overview with core abstractions, a quick-start example, and a pointer to AGENTS.md. Closes apache#1394 (PR 1a) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Clarify audience of root vs package AGENTS.md The root AGENTS.md (symlinked as CLAUDE.md) is for contributors working on the project. Add a pointer to python/datafusion/AGENTS.md which is the user-facing DataFrame API guide shipped with the package. Also add the Apache license header to the package AGENTS.md. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Add PR template and pre-commit check guidance to AGENTS.md Document that all PRs must follow .github/pull_request_template.md and that pre-commit hooks must pass before committing. List all configured hooks (actionlint, ruff, ruff-format, cargo fmt, cargo clippy, codespell, uv-lock) and the command to run them manually. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Remove duplicated hook list from AGENTS.md Let the hooks be discoverable from .pre-commit-config.yaml rather than maintaining a separate list that can drift. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Fix AGENTS.md: Arrow C Data Interface, aggregate filter, fluent example - Clarify that DataFusion works with any Arrow C Data Interface implementation, not just PyArrow. - Show the filter keyword argument on aggregate functions (the idiomatic HAVING equivalent) instead of the post-aggregate .filter() pattern. - Update the SQL reference table to show FILTER (WHERE ...) syntax. - Remove the now-incorrect "Aggregate then filter for HAVING" pitfall. - Add .collect() to the fluent chaining example so the result is clearly materialized. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Update agents file after working through the first tpc-h query using only the text description * Add feedback from working through each of the TPC-H queries * Address Copilot review feedback on AGENTS.md - Wrap CASE/WHEN method-chain examples in parentheses and assign to a variable so they are valid Python as shown (Copilot #1, apache#2). - Fix INTERSECT/EXCEPT mapping: the default distinct=False corresponds to INTERSECT ALL / EXCEPT ALL, not the distinct forms. Updated both the Set Operations section and the SQL reference table to show both the ALL and distinct variants (Copilot apache#4). - Change write_parquet / write_csv / write_json examples to file-style paths (output.parquet, etc.) to match the convention used in existing tests and examples. Note that a directory path is also valid for partitioned output (Copilot apache#5). Verified INTERSECT/EXCEPT semantics with a script: df1.intersect(df2) -> [1, 1, 2] (= INTERSECT ALL) df1.intersect(df2, distinct=True) -> [1, 2] (= INTERSECT) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Use short-form comparisons in AGENTS.md examples Drop lit() on the RHS of comparison operators since Expr auto-wraps raw Python values, matching the style the guide recommends (Copilot apache#3, apache#6). Updates examples in the Aggregation, CASE/WHEN, SQL reference table, Common Pitfalls, Fluent Chaining, and Variables-as-CTEs sections, plus the __init__.py quick-start snippet. Prose explanations of the rule (which cite the long form as the thing to avoid) are left unchanged. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Move user guide from python/datafusion/AGENTS.md to SKILL.md The in-wheel AGENTS.md was not a real distribution channel -- no shipping agent walks site-packages for AGENTS.md files. Moving to SKILL.md at the repo root, with YAML frontmatter, lets the skill ecosystems (npx skills, Claude Code plugin marketplaces, community aggregators) discover it. Update the pointers in the contributor AGENTS.md and the __init__.py module docstring accordingly. The docstring now references the GitHub URL since the file no longer ships with the wheel. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Address review feedback: doctest, streaming, date/timestamp - Convert the __init__.py quick-start block to doctest format so it is picked up by `pytest --doctest-modules` (already the project default), preventing silent rot. - Extract streaming into its own SKILL.md subsection with guidance on when to prefer execute_stream() over collect(), sync and async iteration, and execute_stream_partitioned() for per-partition streams. - Generalize the date-arithmetic rule from Date32 to both Date32 and Date64 (both reject Duration at any precision, both accept month_day_nano_interval), and note that Timestamp columns differ and do accept Duration. - Document the PyArrow-inherited type mapping returned by to_pydict()/to_pylist(), including the nanosecond fallback to pandas.Timestamp / pandas.Timedelta and the to_pandas() footgun where date columns come back as an object dtype. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Distinguish user guide from agent reference in module docstring The docstring pointed readers at SKILL.md as a "comprehensive guide," but SKILL.md is written in a dense, skill-oriented format for agents — humans are better served by the online user guide. Put the online docs first as the primary reference and label the SKILL.md link as the agent reference. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

nuno-faria and others added 14 commits November 29, 2025 19:36

Upgrade to Datafusion 51

6bb9678

Fix clippy

8f65f42

Refactor test_arrow_c_stream_interrupted to handle exceptions in a se…

b991f77

…parate thread

Refactor signal checking in future collection to simplify error handling

0fa8178

Simplify KeyboardInterrupt check in test_arrow_c_stream_interrupted

1297e2c

rm test_record_batch_reader_interrupt_exits_quickly

e8842b1

Refactor test_arrow_c_stream_interrupted to improve exception handlin…

bf8085e

…g and readability

Improve exception handling in test_arrow_c_stream_interrupted to chec…

b10ff3a

…k for KeyboardInterrupt

Add comment - handle KeyboardInterrupt more effectively

3f8e6d9

Remove unnecessary stream cancellation on error in PartitionedDataFra…

23999c8

…meStreamReader

Add jupyter notebook for test

784929d

Revert "Add jupyter notebook for test"

5ec0c78

This reverts commit 784929d.

Simplify error handling in PartitionedDataFrameStreamReader

506091e

kosiew mentioned this pull request Nov 29, 2025

Upgrade to Datafusion 51 apache/datafusion-python#1311

Merged

Merge branch 'datafusion_51' into pr-1311

5615d83

nuno-faria merged commit ab63c6c into nuno-faria:datafusion_51 Nov 29, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Improve KeyboardInterrupt handling during Arrow C stream reads in Python bindings#1

Improve KeyboardInterrupt handling during Arrow C stream reads in Python bindings#1
nuno-faria merged 15 commits into
nuno-faria:datafusion_51from
kosiew:pr-1311

kosiew commented Nov 29, 2025 •

edited

Loading

Uh oh!

kosiew commented Nov 29, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Conversation

kosiew commented Nov 29, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Rationale for this change

What changes are included in this PR?

Are these changes tested?

Test in Jupyter Notebook

Are there any user-facing changes?

Uh oh!

kosiew commented Nov 29, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

kosiew commented Nov 29, 2025 •

edited

Loading