WorkFlow - stuck immediately after the Start event

0
Hi everyone,We are experiencing an unusual issue with the Workflow module. Sometimes a workflow is triggered successfully and the workflow instance is created, but it remains stuck at the Start activity and does not proceed to the next step even after the Start we have a parellel split activity.The issue is intermittent:Some workflow instances run successfully from start to end.Some workflow instances get stuck immediately after the Start event.Has anyone experienced workflows getting stuck at the Start activity after a Workflow module upgrade or Mendix upgrade?Any guidance or similar experiences would be greatly appreciated.Thanks!Mendix version: 11.12.4Worflow Module Version : 4.10.0
asked
3 answers
1

Hi,

I haven't seen a known issue listed for this exact combination, but "stuck at Start" usually means one of a few things. The fact that it's intermittent narrows it down. I'd go through these in order:

1. Check the real state of a stuck instance

The workflow diagram showing it at Start doesn't always mean the engine thinks it's there. Look at the System.Workflow object (via the Workflow Admin pages in WorkflowCommons or a quick overview page) and check:

  • State: InProgress, Failed, Incompatible or Paused
  • Reason: populated if something failed
  • The activity records: whether the parallel split or any activity after it was ever created

If the state is Failed, the problem is in one of the first activities after the split, not the Start event itself.

2. Turn up logging

Set the WorkflowEngine log node to Debug or Trace, then reproduce. Errors in the activities right after the split, such as a Call microflow, a user task's target users microflow/XPath, or a workflow or user task event handler, will show up there. Because the split starts all paths together, one failing path is enough to block or fail the whole instance.

3. Look at how the workflow is triggered

The Call workflow activity creates the instance in the current transaction. The engine only starts executing it after that transaction commits, asynchronously. Intermittent behaviour often comes from:

  • The context object being created or changed in the same long-running microflow, or in a before/after-commit event handler
  • Many workflows started in a loop or batch at once
  • The calling microflow failing or rolling back after the Call workflow activity in some cases

A simple test is to commit the context object first and start the workflow as the last step of the microflow, or from a separate, short microflow.

4. Contact Mendix support (Ticket)
If none of the above explains it, it may be a platform bug, so a support ticket is worth it. Attach the WorkflowEngine logs and the details of a stuck instance. It also helps a lot to share the project so they can reproduce the issue. Export it as a project package (File > Export Project Package, .mpk) rather than sending just the .mpr, since the .mpr alone doesn't include Java actions, widgets and other resources. Even better is a small test project that reproduces the issue with only the affected workflow, so you don't have to share the full app or any sensitive data.


Hope it helps!
Fjordi

answered
1

Hi Guna,


Since the workflow instance is created successfully but some instances remain at the Start activity, I would first check the workflow instance itself rather than the Workflow Commons UI.

The Start activity is normally just the entry point of the workflow. If the workflow remains in In Progress immediately after starting, something is preventing the Workflow Engine from advancing the instance to the next activity.


I would check the following first:

  1. Open the workflow instance in the Workflow Admin page and check its current state and activity.
  2. Check the Runtime log at the exact time the affected workflow is started. If the workflow-initiated microflow throws an exception, the workflow should normally move to Failed, with the reason available on the workflow instance.
  3. Compare a workflow instance that completes successfully with one that gets stuck. Since the problem is intermittent, compare the workflow context data and the first activity after Start.
  4. Check whether there is any parallel execution or concurrency around the workflow context. The Workflow Engine locks the System.Workflow record while executing a workflow instance, so concurrent changes to the same workflow instance are not allowed.
  5. If this started after changing/upgrading the workflow definition, check the Workflow Versioning / Compatibility information. Changes to an already-running workflow can leave instances incompatible in certain cases, particularly when the execution path has changed. Mendix provides the Workflow Admin functionality to continue, restart, jump, or abort affected instances where applicable.


One thing I would also check is the version combination. Workflow Commons 4.10.0 targets Mendix 11.11.0, while your application is running on 11.12.4. That does not by itself prove that Workflow Commons is causing the issue, but I would make sure all Marketplace modules are on versions supported for the Studio Pro version you are using.


If the issue is reproducible, I would temporarily enable more detailed logging around the workflow execution and capture the log entries from the moment the workflow is triggered. That should tell you whether the engine is waiting, failing, or repeatedly attempting to advance the workflow.


I would not manually update the workflow state or database records to move these instances past the Start activity. Use the Workflow Admin actions such as Retry/Restart/Continue where they are applicable.

Since only some instances are affected, the most useful next step would be to compare one successful and one stuck instance and check their workflow state, current activity, context data, and Runtime log entries.


Hope this helps.

answered
1

We've run into a similar issue when workflows were started inside a long-running transaction. The workflow instance was created, but execution didn't continue until the transaction committed. In some cases the calling microflow later rolled back, which left us with workflow instances appearing to be stuck at Start.

What resolved it for us was moving the "Call workflow" activity to the end of the process and ensuring the context object was already committed before starting the workflow.

Since your issue is intermittent, I'd specifically compare how the workflow is being initiated in successful versus stuck cases. If the context object is being modified by events, task queues, or other concurrent processes immediately after workflow creation, that can be worth investigating.


answered