Hi Pavan,
I haven't worked specifically with DocLoader, but I have implemented similar architectures where documents could originate from different repositories and then be processed in Mendix.
My recommendation would be to separate the solution into two layers:
Define a common interface in Mendix:
GetDocument(SourceType, DocumentReference)
↓
Returns FileDocument
For example:
This keeps the rest of your processing independent of where the file originated.
Once the file is stored as a Mendix System.FileDocument, process it using a common flow:
FileDocument
↓
Determine file type
↓
PDF → Extract Text
DOCX → Extract Text
↓
Normalized Text
↓
AI / Business Processing
The first thing I would determine is whether:
https://.../docloader.asp?docid=XXXX
simply renders a web page or ultimately redirects to a downloadable file.
In practice I would:
Many legacy document repositories expose:
even when the user only sees a browser page.
If such an endpoint exists, Mendix can usually call it using:
and save the response into a FileDocument.
Then I'd discuss with the DocLoader administrators before attempting any workaround.
Preferred order:
I would avoid browser automation unless absolutely necessary.
For scalability, I would standardize both formats into plain text before applying business rules or AI extraction.
Typical flow:
PDF/DOCX
↓
Text Extraction
↓
JSON Structure
↓
Business Validation
↓
AI Processing
This allows the same extraction logic regardless of source system.
User clicks Deep Link
↓
Mendix Microflow
↓
Determine Source
├── SharePoint
│ ↓
│ Download File
│
└── DocLoader
↓
Retrieve File
↓
Store as FileDocument
↓
Extract Text
↓
AI / Automation
My biggest recommendation would be: treat SharePoint and DocLoader only as document sources and normalize everything into a Mendix FileDocument as early as possible. Once you do that, your extraction, AI processing, and downstream automation become source-independent and much easier to maintain.