Retrieving Documents from SharePoint and DocLoader for Automated Processing

0
Hi Team,I am working on a requirement where users receive an email containing a table similar to the one below:S.NoLink to DocumentRemarkAI1SharePoint Document LinkTestBot Link (Deep Link to trigger Microflow)2DocLoader Document LinkTestBot Link (Deep Link to trigger Microflow)When a user clicks the AI link, a Mendix microflow is triggered. The objective is:Retrieve the document referenced in the corresponding row.Process the document content of type PDF and WordExtract data from the document for further automation.Current StatusI am able to download documents that are stored in SharePoint using the existing SharePoint integration.For documents stored in DocLoader, the link looks similar to the following:https://eos-web.erlm.abc.de/docloader/docloader.asp?docid=EG524KL-HJ525-86UJ-R8H4-J5K6I2E11QuestionsHas anyone worked with DocLoader integrations before?Is there a supported way to programmatically retrieve the actual file content from a DocLoader URL (or from the corresponding docid) rather than simply opening the link in a browser?Does DocLoader expose any API or download endpoint that can be called from Mendix/Java to obtain the document as a file?Once the file is obtained, what would be the recommended approach to handle different document formats (PDF, word docx) and extract their content in a scalable way?Are there any best practices for building a generic document-processing flow where the source document may come from either SharePoint or DocLoader?Any guidance, architecture recommendations, or previous implementation experience would be greatly appreciated.Thanks in advance.Note : I am using mendix 9.24 version
asked
1 answers
0

Hi Pavan,

I haven't worked specifically with DocLoader, but I have implemented similar architectures where documents could originate from different repositories and then be processed in Mendix.

My recommendation would be to separate the solution into two layers:

1. Document Retrieval Layer

Define a common interface in Mendix:

GetDocument(SourceType, DocumentReference)
     ↓
Returns FileDocument

For example:

  • SharePoint → Use the SharePoint Connector to download the file.
  • DocLoader → Implement a dedicated retrieval microflow/Java action.

This keeps the rest of your processing independent of where the file originated.

2. Document Processing Layer

Once the file is stored as a Mendix System.FileDocument, process it using a common flow:

FileDocument
     ↓
Determine file type
     ↓
PDF → Extract Text
DOCX → Extract Text
     ↓
Normalized Text
     ↓
AI / Business Processing

Regarding DocLoader

The first thing I would determine is whether:

https://.../docloader.asp?docid=XXXX

simply renders a web page or ultimately redirects to a downloadable file.

In practice I would:

  1. Open browser developer tools.
  2. Access the DocLoader link.
  3. Inspect the Network tab.
  4. Check whether an API call or file download endpoint is used behind the scenes.

Many legacy document repositories expose:

  • a download URL
  • a document ID endpoint
  • a content service

even when the user only sees a browser page.

If such an endpoint exists, Mendix can usually call it using:

  • Call REST
  • Call HTTP
  • Custom Java action

and save the response into a FileDocument.

If DocLoader Has No API

Then I'd discuss with the DocLoader administrators before attempting any workaround.

Preferred order:

  1. Official REST/SOAP API
  2. Download endpoint using document ID
  3. Service account integration
  4. Browser automation (last resort)

I would avoid browser automation unless absolutely necessary.

PDF and Word Processing

For scalability, I would standardize both formats into plain text before applying business rules or AI extraction.

Typical flow:

PDF/DOCX
     ↓
Text Extraction
     ↓
JSON Structure
     ↓
Business Validation
     ↓
AI Processing

This allows the same extraction logic regardless of source system.

Recommended Overall Architecture

User clicks Deep Link
       ↓
Mendix Microflow
       ↓
Determine Source
       ├── SharePoint
       │      ↓
       │   Download File
       │
       └── DocLoader
              ↓
           Retrieve File
              ↓
       Store as FileDocument
              ↓
       Extract Text
              ↓
       AI / Automation

My biggest recommendation would be: treat SharePoint and DocLoader only as document sources and normalize everything into a Mendix FileDocument as early as possible. Once you do that, your extraction, AI processing, and downstream automation become source-independent and much easier to maintain.


answered