Repository navigation
Binary (windows?) wheels for 3.15 are probably incompatible with newly-released 3.15b4; ~perhaps a 3.15b4 incompatibility in general~ #263
Description
Activity
Are you set up to run your tests under a debugger on those platforms? Typically that kind of issue crops up when an application's cffi-allocated variable is not held in the interpreter and the GC collects the buffer while something else is accessing that buffer. It would help if you could cut the failure down to a small reproducer.
I can reproduce one of the smaller crashes I saw locally on macOS (I have no access to Windows). Using the pre-built PyPI wheel on 3.15b4,
lldb python -m gevent.tests.test__corecrashes like so:* thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) frame #0: 0x00000001013709b4 Python`take_gil + 60 Python`take_gil: -> 0x1013709b4 <+60>: ldr x22, [x20, #0x10] 0x1013709b8 <+64>: add x0, x22, #0x50 0x1013709bc <+68>: bl 0x1014c4da4 ; symbol stub for: pthread_mutex_lock 0x1013709c0 <+72>: cbnz w0, 0x101370c6c ; <+756> Target 0: (Python) stopped. (lldb) bt * thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) * frame #0: 0x00000001013709b4 Python`take_gil + 60 frame #1: 0x0000000101371074 Python`PyEval_RestoreThread + 68 frame #2: 0x0000000105d9402c _corecffi.abi3.so`_cffi_f_uv_run + 348 frame #3: 0x00000001011dc358 Python`cfunction_call + 108 frame #4: 0x00000001012f9854 Python`_TAIL_CALL_CALL + 1828 frame #5: 0x00000001012ed554 Python`_PyEval_Vector + 692 ...If I manually build CFFI from source, the test passes fine.
I can boil that test case down to a simpler script:
import gevent gevent.config.loop = 'libuv-cffi' gevent.sleep(1)
This too crashes in the same way with the pre-built CFFI wheel, and works correctly with a newly locally built wheel:
# lldb -- python cffi-crash.py (lldb) target create "python" Current executable set to '/Users/jmadden/Projects/VirtualEnvs/tmp-bfa8feaa5f3c1e0/bin/python' (arm64). (lldb) settings set -- target.run-args "cffi-crash.py" (lldb) run Process 15421 launched: '/Users/jmadden/Projects/VirtualEnvs/tmp-bfa8feaa5f3c1e0/bin/python' (arm64) Process 15421 stopped * thread #2, stop reason = exec frame #0: 0x00000001000189c0 dyld`_dyld_start dyld`_dyld_start: -> 0x1000189c0 <+0>: mov x0, sp 0x1000189c4 <+4>: and sp, x0, #0xfffffffffffffff0 0x1000189c8 <+8>: mov x29, #0x0 ; =0 0x1000189cc <+12>: mov x30, #0x0 ; =0 Target 0: (Python) stopped. (lldb) c Process 15421 resuming Process 15421 stopped * thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) frame #0: 0x00000001013709b4 Python`take_gil + 60 Python`take_gil: -> 0x1013709b4 <+60>: ldr x22, [x20, #0x10] 0x1013709b8 <+64>: add x0, x22, #0x50 0x1013709bc <+68>: bl 0x1014c4da4 ; symbol stub for: pthread_mutex_lock 0x1013709c0 <+72>: cbnz w0, 0x101370c6c ; <+756> Target 0: (Python) stopped. (lldb) bt * thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) * frame #0: 0x00000001013709b4 Python`take_gil + 60 frame #1: 0x0000000101371074 Python`PyEval_RestoreThread + 68 frame #2: 0x0000000105caf04c _corecffi.abi3.so`_cffi_f_uv_run + 608 frame #3: 0x00000001011dc358 Python`cfunction_call + 108 frame #4: 0x00000001012f9854 Python`_TAIL_CALL_CALL + 1828 frame #5: 0x00000001012ed554 Python`_PyEval_Vector + 692 frame #6: 0x000000010115640c Python`_PyObject_VectorcallPrepend + 368 frame #7: 0x000000010002fcdc _greenlet.cpython-315-darwin.so`greenlet::UserGreenlet::inner_bootstrap(_greenlet*, _object*) + 204 frame #8: 0x000000010002f5f4 _greenlet.cpython-315-darwin.so`greenlet::UserGreenlet::g_initialstub(void*) + 1200 frame #9: 0x000000010002e258 _greenlet.cpython-315-darwin.so`greenlet::UserGreenlet::g_switch() + 232 frame #10: 0x0000000100032898 _greenlet.cpython-315-darwin.so`green_switch(_greenlet*, _object*, _object*) + 292 frame #11: 0x0000000105a13f0c _gevent_c_greenlet_primitives.cpython-315-darwin.so`__pyx_f_6gevent_29_gevent_c_greenlet_primitives__greenlet_switch + 80 frame #12: 0x0000000105a105a8 _gevent_c_greenlet_primitives.cpython-315-darwin.so`__pyx_f_6gevent_29_gevent_c_greenlet_primitives_25SwitchOutGreenletWithLoop_switch + 2104 frame #13: 0x0000000105a6ad30 _gevent_c_waiter.cpython-315-darwin.so`__pyx_f_6gevent_16_gevent_c_waiter_6Waiter_get + 3280 frame #14: 0x0000000105ae63f4 _gevent_c_hub_primitives.cpython-315-darwin.so`__pyx_f_6gevent_24_gevent_c_hub_primitives_22WaitOperationsGreenlet_wait + 1780 frame #15: 0x0000000105aeb538 _gevent_c_hub_primitives.cpython-315-darwin.so`__pyx_pf_6gevent_24_gevent_c_hub_primitives_22WaitOperationsGreenlet_wait + 64 frame #16: 0x0000000105aead90 _gevent_c_hub_primitives.cpython-315-darwin.so`__pyx_pw_6gevent_24_gevent_c_hub_primitives_22WaitOperationsGreenlet_1wait + 40 frame #17: 0x0000000105af398c _gevent_c_hub_primitives.cpython-315-darwin.so`__Pyx_CyFunction_Vectorcall_O + 240
I'm happy to try to provide any other helpful information.
- changed the title
[-]Binary (windows?) wheels for 3.15 are probably incompatible with newly-released 3.15b4; perhaps a 3.15b4 incompatibility in general[/-][+]Binary (windows?) wheels for 3.15 are probably incompatible with newly-released 3.15b4; ~perhaps a 3.15b4 incompatibility in general~[/+]on Jul 22, 2026 I can take greenlet and most of gevent out of the equation altogether; once again, this crashes with the pre-built wheel and works with the built-from-source wheel.
import gevent gevent.config.loop = 'libuv-cffi' loop = gevent.get_hub().loop # uses the libuv CFFI interface to libuv io = loop.io(1, 1) io.start(lambda *_: None) loop.run() # libuv.uv_run
* thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) frame #0: 0x00000001013709b4 Python`take_gil + 60 Python`take_gil: -> 0x1013709b4 <+60>: ldr x22, [x20, #0x10] 0x1013709b8 <+64>: add x0, x22, #0x50 0x1013709bc <+68>: bl 0x1014c4da4 ; symbol stub for: pthread_mutex_lock 0x1013709c0 <+72>: cbnz w0, 0x101370c6c ; <+756> Target 0: (Python) stopped. (lldb) bt * thread #2, queue = 'com.apple.main-thread', stop reason = EXC_BAD_ACCESS (code=1, address=0x10) * frame #0: 0x00000001013709b4 Python`take_gil + 60 frame #1: 0x0000000101371074 Python`PyEval_RestoreThread + 68 frame #2: 0x0000000105beb04c _corecffi.abi3.so`_cffi_f_uv_run + 608 frame #3: 0x00000001011dc358 Python`cfunction_call + 108 frame #4: 0x00000001012f9854 Python`_TAIL_CALL_CALL + 1828 frame #5: 0x00000001012ed554 Python`_PyEval_Vector + 692 frame #6: 0x00000001012ed1fc Python`PyEval_EvalCode + 160 frame #7: 0x00000001013f77b4 Python`run_mod + 360 frame #8: 0x00000001013f74c8 Python`_PyRun_File + 156 frame #9: 0x00000001013f6f70 Python`_PyRun_SimpleFile + 232 frame #10: 0x00000001013f6c04 Python`_PyRun_AnyFile + 80 frame #11: 0x000000010142aaa8 Python`pymain_run_file_obj + 160 frame #12: 0x000000010142a18c Python`pymain_run_file + 72 frame #13: 0x00000001014299b0 Python`Py_RunMain + 1644 frame #14: 0x000000010142ad8c Python`pymain_main + 488 frame #15: 0x000000010142af2c Python`Py_BytesMain + 44 frame #16: 0x0000000186b6be00 dyld`start + 6992Interesting that the crash is around the GIL handling. I would expect there to be some difference in the include files that could influence this: some macro or struct change. But
$ git diff v3.15.0b4 v3.15.0b3 Include/did not point out anything obvious except for an additional field in_Py_DebugOffsets, an internal struct that I would be surprised that cffi touches.Ahh,
_PyRun_AnyFilechanged, it now returns aPyObject*and not anint. But that does not come from cffi, it comes from somewhere else. So perhaps by using a cffi source build, something else in the build changes?The
PyThreadState(struct _ts) also changed size, it grewuintptr_t last_profiled_frame_seqin 3.15b4 (that's the change you noticed inPy_DebugOffsetsas well). But that should be an opaque object, always handled asPyThreadState*, never directly allocated orsizeof()or anything like that, and I don't see anywhere that cffi is reading or writing directly toPyThreadStatefields. But because that's the argument toPyEval_RestoreThreadI'm suspicious of that change.There was an ABI break from b3 to b4: python/cpython#152448.
Fixing this probably requires uploading new CFFI wheels for Python 3.15, which might necessitate doing a new release.
- added a commit that references this issue
on Jul 29, 2026 @nitzmahone do you think you'll have time soonish to coordinate a release to help with this? I know @mattip will be away for a while too.
If everyone's reasonably confident it's only the ABI change and that just re-spinning the wheels against b4 is sufficient, I can probably do that this morning. I might try to tweak the
requires-pythonconstraint to block 3.15.0a0-3.15.0b3, since I'm guessing the offending ABI change could be equally fatal in reverse.I'll ask about that on the core dev discord
Reacted by Matt DavisHere's a pure-CFFI reproducer, which crashes for me on 3.15.0b4 using the CFFI wheel on PyPI
import sys import tempfile import cffi ffi = cffi.FFI() ffi.cdef(""" extern "Python" void event_cb(void); void run_one(void); """) ffi.set_source("_repro263", """ static void event_cb(void); static void run_one(void) { event_cb(); } """) tmpdir = tempfile.mkdtemp() ffi.compile(tmpdir=tmpdir) sys.path.insert(0, tmpdir) from _repro263 import ffi, lib @ffi.def_extern() def event_cb(): print("in callback") lib.run_one() print("no crash")@nitzmahone I think the answer is "no, barring a late-breaking bug that justifies an ABI break". So, hopefully not?
I still think it'd be worth doing a new bugfix release for this to help unblock downstream projects that depend on CFFI who want to set up Python 3.15 testing. But I also understand your time and attention is limited and I very much appreciate you spending time on this.
I still think it'd be worth doing a new bugfix release
Oh, no argument there- I'll be pushing rebuilt wheels for b4+ support regardless. Just wanted to make sure there was reasonable confidence that no code changes are required before I do that.
Thanks for the small standalone repro BTW!
I've hit similar situations before on other projects where it's only a packaging issue with literally zero code change, but it's been awhile. In former lives, I'd just manually publish new wheels as
2.1.0.post1or something, but with all the supply-chain security bots already looking sideways at us for not using attested publishing (hoping to wait for GHA's lockfile feature to land first), not having a matching upstream release + tag would probably cause more trouble than just shipping a full 2.1.1: "same as 2.1.0, now with less stack corruption in 3.15 wheels!" release 🤷 . That's probably what I'll end up doing.I think I actually have a way to use Python C API calls alone to avoid the fragile direct struct access entirely, PR incoming.
See #269. There's still one direct PyThreadState access left after that PR but it can't be deleted yet and it's not in a public API so it's less of a problem.
Reacted by Matt DavisI'll do some CI spot-checks against older Pythons (looks like "oldest-supported" == 3.9 fell out of the PR matrix at some point, which was unintentional), but assuming all's good, I'll merge that and do a 2.1.1 release. I'm always leery about doing late-in-week releases, so might hold up the actual build and push until Monday morning since I'd be unavailable to address any issues until then anyway.
That's fine with me. This is worth fixing but given that 3.15 is still a prerelease it's not a fire or anything like that.
Reacted by Matt Davis- added a commit that references this issue
on Aug 1, 2026 I ran into this on Windows, macOS, Ubuntu. Also, it appears cffi 2.0.0 is immune! Surprising. (For now I'll just stick with cffi 2.0.0...)
fixed by #269 in 2.1.1
Over in this cython bug I documented a case where binary wheels that worked in 3.15b3 failed in 3.15b4, which came out this weekend.
I can't prove it, but I'm pretty sure CFFI's wheels are in the same situation. I believe this because, without any related code changes on gevent's part, our Windows builds that worked correctly on 3.15b3 suddenly started segfaulting in nearly every test on 3.15b4.
There is some additional information. Even, after I started building CFFI from scratch on mac/linux ((EDIT: I think I must still have been using the pre-built wheel) I was seeing one crash on 3.15 that appears related to CFFI. That's the only test I had that uses CFFI (the Windows tests all use CFFI), and when I disabled that test, the crashes went away, so I think we can rule out gevent or 3.15b4 itself.pip install --no-binary=:all: cffi),