Skip to content

[Feature Request]: PyMuPDF4LLM for building KnowledgeBase (with lightrag) #589

Description

@VarLad

Do you need to file a feature request?

  • I have searched the existing feature request and this feature request is not already filed.
  • I believe this is a legitimate feature request, not just a question or bug.

Feature Request Description

Provide a lightweight alternative to mineru and docling for lightrag (that can extract images) like PyMuPDF4LLM.
This would enable DeepTutor (and lightrag) to easily run on lower-end devices

Related Module

Knowledge Base Management

Use Case

Hi, I was looking at the options lightrag has, and almost all the options (that can work with pdf documents and extract images), i.e., mineru and docling pull a CUDA dependency. I'm running DeepTutor on a device similar to a Raspberry Pi.
It would be nice, if something lightweight like PyMuPDF4LLM could be used. Link here: https://github.com/pymupdf/pymupdf4llm
To me it seems that this solution is quite efficient, although a bit time taking.

Do tell me if this issue makes more sense to open in the lightrag repo.

Additional Context

No response

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions