Skip to content

Fix suspend resolvers hanging when awaiting a data loader - #831

Merged
oryan-block merged 3 commits into
masterfrom
bugfix/419
Oct 4, 2026
Merged

oryan-block merged 3 commits into
masterfrom
bugfix/419

Conversation

@oryan-block

@oryan-block oryan-block commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes #419

Checklist

  • Pull requests follows the contribution guide
  • New or modified functionality is covered by tests

Description

A suspend resolver that calls DataLoader#load or loadMany and then awaits the result hangs forever. This still happens on master (graphql-java 26.1) with default options, for root and nested fields. MethodFieldResolverDataFetcher#get started suspend resolvers with the default coroutine start, which hands the body to the configured dispatcher (Dispatchers.Default unless configured otherwise) and returns the future straight away. graphql-java dispatches a level's DataLoaders once that level's data fetchers have returned, so the load usually got queued after the dispatch had already run, and the await waited for a batch that never ran. It's basically the supplyAsync example the graphql-java batching docs tell you not to do, as someone pointed out on the issue.

The coroutine now starts with CoroutineStart.UNDISPATCHED, so the resolver runs on the calling thread until its first suspension. Loads are queued before the data fetcher returns and the rest of the coroutine still resumes on the configured context. An undispatched coroutine runs its body even if its context is already cancelled, which a default start never does, so the block calls ensureActive() first. That way a resolver still isn't invoked when, for example, the Job passed through SchemaParserOptions.coroutineContext is cancelled. Calling dispatch() by hand, like the workaround on the issue does, shouldn't be needed anymore for this case.

SuspendFunctionDataLoaderTest covers a root suspend field awaiting loadMany and a nested suspend field awaiting load. Each runs 20 executions with a 5s timeout so a hang fails the test instead of blocking the build. Both time out on master. There's also a new test in MethodFieldResolverDataFetcherTest for the cancelled context.

This only covers loads issued before the resolver first suspends. A load issued after a real suspension, like a second load that depends on the first or a load after withContext(Dispatchers.IO), still hangs with default options, on master and with this change. A plain fetcher that chains CompletableFuture loads hangs the same way, so I didn't try to work around it here. For those, graphql-java's DataLoaderDispatchingContextKeys.ENABLE_DATA_LOADER_CHAINING or ENABLE_DATA_LOADER_EXHAUSTED_DISPATCHING work together with this change. Chaining only helps loaders that come from env.getDataLoader, not ones held on a custom context object. Dispatchers.Unconfined isn't a full workaround either. It still hangs on master for suspend fields nested under another suspend field. Subscriptions are left alone.

Behaviour change: code in a suspend resolver up to its first suspension, including the resolver method call itself, now runs on the graphql-java calling thread instead of the dispatcher from SchemaParserOptions.coroutineContext / coroutineContextProvider. After the first suspension it resumes on the configured context as before, and the coroutine context itself is unchanged. So a suspend function that does blocking work without ever suspending (e.g. blocking JDBC with no withContext) now blocks the calling thread, same as a non-suspend resolver. Since graphql-java calls sibling data fetchers one after another on that thread, sibling fields backed by such resolvers now run one after another instead of in parallel, so a query with several of them gets slower without any error. Wrapping the blocking call in withContext(Dispatchers.IO) moves it off again and brings the parallelism back. This should go in the release notes. Exceptions thrown before the first suspension still fail the field the same way as before.

🤖 Generated with Claude Code

oryan-block and others added 3 commits October 3, 2026 18:30
Suspend resolver methods were started with the default coroutine start,
which dispatches the body to the configured dispatcher and returns the
future straight away. graphql-java dispatches the DataLoaders of a
level once the data fetchers of that level have returned, so a resolver
that awaited DataLoader.load or loadMany often registered its keys only
after that dispatch had happened and then waited forever for a batch
that never ran.

Start the coroutine with CoroutineStart.UNDISPATCHED so the resolver
runs on the calling thread until its first suspension. Loads are then
queued before the data fetcher returns, as graphql-java expects, and
the rest of the coroutine still resumes on the configured context. An
undispatched coroutine runs even if its context is already cancelled,
so check for that first to keep cancelled resolvers from being invoked.

Loads issued after the resolver has suspended, such as a second load
that depends on the first, still rely on graphql-java's data loader
chaining or exhausted dispatching options, as with plain fetchers.

Fixes #419

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sonarqubecloud

sonarqubecloud Bot commented Oct 4, 2026

Copy link
Copy Markdown

@oryan-block
oryan-block merged commit ee213e5 into master Oct 4, 2026
6 checks passed
@oryan-block
oryan-block deleted the bugfix/419 branch October 4, 2026 17:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Kotlin coroutine suspend fun resolver hangs when awaiting a data loader

1 participant