Load from the Huawei OBS file.
Parse Oracle doc metadata...
Read a file
Load Jupyter notebook (.ipynb) files.
Load from Amazon AWS S3 directory.
Load documents from TiDB.
Load local Airbyte json files.
Load a sitemap and its URLs.
Security Note: This loader can be used to load all URLs specified in a sitemap. If a malicious actor gets access to the sitemap, they could force the server t
Load from TensorFlow Dataset.
Load CHM files using Unstructured.
CHM means Microsoft Compiled HTML Help.
Examples
from langchain_community.document_loaders import UnstructuredCHMLoader
loader = UnstructuredCHMLoade
Microsoft Compiled HTML Help (CHM) Parser.
Load Org-Mode files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document
Load from Hugging Face Hub datasets.
Load Roam files from a directory.
Load Pandas DataFrame.
Load files from Dropbox.
In addition to common files such as text and PDF files, it also supports Dropbox Paper files.
Load pages from OneNote notebooks.
Load from Telegram chat dump.
Load Telegram chat json directory dump.
Load Documents using LLMSherpa.
LLMSherpaFileLoader use LayoutPDFReader, which is part of the LLMSherpa library. This tool is designed to parse PDFs while preserving their layout information, which
Pebblo Safe Loader class is a wrapper around document loaders enabling the data to be scrutinized.
Loader for text data.
Since PebbloSafeLoader is a wrapper around document loaders, this loader is used to load text data directly into Documents.
Load iFixit repair guides, device wikis and answers.
iFixit is the largest, open repair community on the web. The site contains nearly 100k repair manuals, 200k Questions & Answers on 42k devices,
Load DOCX file using docx2txt and chunks at character level.
Defaults to check for local file, but if the file is a web path, it will download it to a temporary file, and use that, then clean up
Load Microsoft Word file using Unstructured.
Works with both .docx and .doc files. You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will
Load from Huawei OBS directory.
Load from Snowflake API.
Each document represents one row of the result. The page_content_columns
are written into the page_content of the document. The metadata_columns
are written into the
Load model information from Hugging Face Hub, including README content.
This loader interfaces with the Hugging Face Models API to fetch and load model metadata and README files. The API allows you
Load elements from a blockchain smart contract.
The supported blockchains are: Near mainnet, Near testnet.
If no BlockchainType is specified, the default is Near mainnet.
The Loader uses the Mintba
Load RTF files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document obje
Load PDF files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document obje
Base Loader class for PDF files.
If the file is a web path, it will download it to a temporary file, use it, then clean up the temporary file after completion.
Load online PDF.
Load and parse a PDF file using 'pypdf' library.
This class provides methods to load and parse PDF documents, supporting various configurations such as handling password-protected files, extrac
Load and parse a PDF file using the pypdfium2 library.
This class provides methods to load and parse PDF documents, supporting various configurations such as handling password-protected files
Load and parse a directory of PDF files using 'pypdf' library.
This class provides methods to load and parse multiple PDF documents in a directory, supporting options for recursive search, handling p
Load and parse a PDF file using 'pdfminer.six' library.
This class provides methods to load and parse PDF documents, supporting various configurations such as handling password-protected files,
Load PDF files as HTML content using PDFMiner.
Load and parse a PDF file using 'PyMuPDF' library.
This class provides methods to load and parse PDF documents, supporting various configurations such as handling password-protected files, extr
Load PDF files using Mathpix service.
Load PDF files using pdfplumber.
Load PDF files from a local file system, HTTP or S3.
To authenticate, the AWS client uses the following methods to automatically load credentials: https://boto3.amazonaws.com/v1/documentation/api/l
DedocPDFLoader document loader integration to load PDF files using dedoc.
The file loader can automatically detect the correctness of a textual layer in the
PDF document.
Note that __init__ me
Load a PDF with Azure Document Intelligence
Document loader utilizing Zerox library: https://github.com/getomni-ai/zerox
Zerox converts PDF document to series of images (page-wise) and uses vision-capable LLM model to generate Markdown represe
Load documents from Yuque.
Load from Open City.
Load Xorbits DataFrame.
Client for lakeFS.
Load from lakeFS.
Load from lakeFS as unstructured data.
Load documents from AWS Athena.
Each document represents one row of the result.
page_content of the document
and none into the metadata of the docLoad documents from Microsoft OneDrive.
Uses SharePointLoader under the hood.
Load from Baidu Cloud BOS file.
Load EPub files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document obj
Load conversations from exported ChatGPT data.
Load webpages with Browserless /content endpoint.
Scrape HTML pages from URLs using a headless instance of the Chromium.
Load HTML asynchronously.
Load CoNLL-U files.
Load files from remote URLs using Unstructured.
Use the unstructured partition function to detect the MIME type and route the file to the appropriate partitioner.
You can run the loader in one of
Load image captions.
By default, the loader utilizes the pre-trained Salesforce BLIP image captioning model. https://huggingface.co/Salesforce/blip-image-captioning-base
Load Notion directory dump.
Load from IUGU.
Load from Azure AI Data.
Load from FaunaDB.
Load MongoDB documents.
WebBaseLoader document loader integration
Load from a directory.
Load records from an ArcGIS FeatureLayer.
Load Quip pages.
Port of https://github.com/quip/quip-api/tree/master/samples/baqup
Load and pars Documents concurrently.
Transcript format to use for the document loader.
Load AssemblyAI audio transcripts.
It uses the AssemblyAI API to transcribe audio files and loads the transcribed text into one or more Documents, depending on the specified format.
To use, you shou
Load AssemblyAI audio transcripts.
It uses the AssemblyAI API to get an existing transcription and loads the transcribed text into one or more Documents, depending on the specified format.
Load TOML files.
It can load a single source file or several files in a single directory.
Load the Airtable tables.
Load College Confidential webpages.
Load Polars DataFrame.
Load geopandas Dataframe.
Generic Document Loader.
A generic document loader that allows combining an arbitrary blob loader with a blob parser.
Examples:
Parse a specific PDF file:
.. code-block:: python
f
Load SurrealDB documents.
Load a query result from Arxiv.
The loader converts the original PDF format into the text.
Load a bibtex file.
Each document represents one entry from the bibtex file.
If a PDF file is present in the file bibtex field, the original PDF
is loaded into the document text. If no such file
Generic Google API Client.
To use, you should have the google_auth_oauthlib,youtube_transcript_api,google
python package installed.
As the google api expects credentials you need to set up a goog
Output formats of transcripts from YoutubeLoader.
Load YouTube video transcripts.
Load all Videos from a YouTube Channel.
To use, you should have the googleapiclient,youtube_transcript_api
python package installed.
As the service needs a google_api_client, you first have to
Load news articles from RSS feeds using Unstructured.
Load Cube semantic layer metadata.
Load from LarkSuite (FeiShu).
Load from LarkSuite (FeiShu) wiki.
Load notes from Joplin.
In order to use this loader, you need to have Joplin running with the Web Clipper enabled (look for "Web Clipper" in the app settings).
To get the access token, you need to
Load from Alibaba Cloud MaxCompute table.
Load Twitter tweets.
Read tweets of the user's Twitter handle.
First you need to go to
https://developer.twitter.com/en/docs/twitter-api /getting-started/getting-access-to-the-twitter-api
to get
Load Datadog logs.
Logs are written into the page_content and into the metadata.
Load documents from Couchbase.
Each document represents one row of the result. The page_content_fields are
written into the page_contentof the document. The metadata_fields are written
into t
Load from Spreedly API.
Load documents by querying database tables supported by SQLAlchemy.
For talking to the database, the document loader uses the SQLDatabase
utility from the LangChain integration toolkit.
Each docum
Load IMSDb webpages.
Load Figma file.
Base class for all loaders that uses O365 Package
Enumerator of the content formats of Confluence page.
Load Confluence pages.
Port of https://llamahub.ai/l/confluence This currently supports username/api_key, Oauth2 login, personal access token or cookies authentication.
Specify a list page_ids and
Load with an Airbyte source connector implemented using the CDK.
A wrapper around the CDK integration.
Load from Hubspot using an Airbyte source connector.
Load from Stripe using an Airbyte source connector.
Load from Typeform using an Airbyte source connector.
Load from Zendesk Support using an Airbyte source connector.
Load from Shopify using an Airbyte source connector.
Load from Salesforce using an Airbyte source connector.
Load from Gong using an Airbyte source connector.
Load ReadTheDocs documentation directory.
Load from a Slack directory dump.
Load AZLyrics webpages.
Load from Kinetica API.
Each document represents one row of the result. The page_content_columns
are written into the page_content of the document. The metadata_columns
are written into the `
Load a PDF with Azure Document Intelligence.
Load Obsidian files from directory.
Document loader for EverNote ENEX export files.
Loads EverNote notebook export files (.enex format) into LangChain Documents.
Extracts plain text content from HTML and preserves note metadata inc
Load Python files, respecting any non-default encoding if specified.
Load Hacker News data.
It loads data from either main page results or the comments page.
Load Markdown files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document
Load weather data with Open Weather Map API.
Reads the forecast & current weather of any location using OpenWeatherMap's free API. Checkout 'https://openweathermap.org/appid' for more on how to gen
File encoding as the NamedTuple.
NeedleLoader is a document loader for managing documents stored in a collection.
Load from SharePoint.
Load from any file type using Nuclia Understanding API.
Load Microsoft PowerPoint files using Unstructured.
Works with both .ppt and .pptx files. You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the docume
Base Loader that uses dedoc (https://dedoc.readthedocs.io).
Loader enables extracting text, tables and attached files from the given file:
* Text can be split by pages, dedoc tree nodes, te
DedocFileLoader document loader integration to load files using dedoc.
The file loader automatically detects the file type (with the correct extension). The list of supported file types is gives at
Load files using dedoc API.
The file loader automatically detects the file type (even with the wrong extension).
By default, the loader makes a call to the locally hosted dedoc API.
More informati
Load .srt (subtitle) files.
Load Diffbot json file.
Load from Tencent Cloud COS directory.
Load PySpark DataFrames.
Column not found error.
Load from a Rockset database.
To use, you should have the rockset python package installed.
Turn a url to llm accessible markdown with Scrapfly.io.
For further details, visit: https://scrapfly.io/docs/sdk/python
Load from DuckDB.
Each document represents one row of the result. The page_content_columns
are written into the page_content of the document. The metadata_columns
are written into the `metada
Load GitBook data.
When load_all_paths=True, the loader parses XML sitemaps an
Load a CSV file into a list of Document objects.
Each document represents one row of the CSV file. Every row is converted into a key/value pair and outputted to a new line in the document's page_
Load CSV files using Unstructured.
Like other Unstructured loaders, UnstructuredCSVLoader can be used in both "single" and "elements" mode. If you use the loader in "elements" mode, the CSV file
Load a Blackboard course.
This loader is not compatible with all Blackboard courses. It is only compatible with courses that use the new Blackboard interface. To use this loader, you must have the
Load from Gutenberg.org.
Load acreom vault from a directory.
Load from Stripe API.
Load XML file using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document objec
Merge documents from a list of loaders
Load from Baidu BOS directory.
Load Facebook Chat messages directory dump.
ModuleName document loader integration
Load TSV files using Unstructured.
Like other Unstructured loaders, UnstructuredTSVLoader can be used in both "single" and "elements" mode. If you use the loader in "elements" mode, the TSV file
Load from Amazon AWS S3 file.
Load PNG and JPG files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Doc
Load a JSON file using a jq schema.
Abstract base class for all evaluators.
Each evaluator should take a page, a browser instance, and a response object, process the page as necessary, and return the resulting text.
Evaluate the page HTML content using the unstructured library.
Load HTML pages with Playwright and parse with Unstructured.
This is useful for loading pages that require javascript to render.
Load HTML using 2markdown API.
Enumerator of the supported blockchains.
Load elements from a blockchain smart contract.
See supported blockchains here: https://python.langchain.com/v0.2/api_reference/community/document_loaders/langchain_community.document_loaders.blockch
Load from Docusaurus Documentation.
It leverages the SitemapLoader to loop through the generated pages of a Docusaurus Documentation website and extracts the content by looking for specific HTML tags
Load a file from Microsoft OneDrive.
Load MediaWiki dump from an XML file.
Load RST files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document obje
Load the Mastodon 'toots'.
Recursively load all child links from a root URL.
Security Note: This loader is a crawler that will start crawling at a given URL and then expand to crawl child links recursively.
We
Load text file.
Parse MHTML files with BeautifulSoup.
Load Git repository files.
The Repository can be local on disk available at repo_path,
or remote at clone_url that will be cloned to repo_path.
Currently, supports only text files.
Each docu
Load from Wikipedia.
The hard limit on the length of the query is 300 for now.
Each wiki page represents one Document.
Load OpenOffice ODT files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Do
FireCrawlLoader document loader integration
Load news articles from URLs using Unstructured.
Load Reddit posts.
Read posts on a subreddit. First, you need to go to https://www.reddit.com/prefs/apps/ and create your application
Load HTML pages with Selenium and parse with Unstructured.
This is useful for loading pages that require javascript to render.
Load cards from a Trello board.
Load from Modern Treasury.
Load from the PubMed biomedical library.
Base Loader that uses Unstructured.
Load transactions from Ethereum mainnet.
The Loader use Etherscan API to interact with Ethereum mainnet.
ETHERSCAN_API_KEY environment variable must be set use this loader.
Load HTML files using Unstructured.
You can run the loader in one of two modes: "single" and "elements". If you use "single" mode, the document will be returned as a single langchain Document obj
Load WhatsApp messages text file.
Load email files using Unstructured.
Works with both .eml and .msg files. You can process attachments in addition to the e-mail message itself by passing process_attachments=True into the construct
Loads Outlook Message files using extract_msg.
https://github.com/TeamMsgExtractor/msg-extractor
Load table schemas from AWS Glue.
This loader fetches the schema of each table within a specified AWS Glue database. The schema details include column names and their data types, similar to pandas dt
Load content from RSpace notebooks, folders, documents or PDF Gallery files.
Map RSpace document <-> Langchain Document in 1-1. PDFs are imported using PyPDF.
Requirements are rspace_client (`pip in
Load with Brave Search engine.
Load from Notion DB.
Reads content from pages within a Notion Database. Args: integration_token (str): Notion integration token. database_id (str): Notion database id. request_timeout_s
Load from Tencent Cloud COS file.
Load Discord chat logs.
Load from Psychic.dev.
Turn an url to LLM accessible markdown with ScrapingAnt.
For further details, visit: https://docs.scrapingant.com/python-client
Load GitHub repository Issues.
Load issues of a GitHub repository.
Load GitHub File
Load fetching transcripts from BiliBili videos.
Load pre-rendered web pages using a headless browser hosted on Browserbase.
Depends on browserbase and playwright packages.
Get your API key from https://browserbase.com
Load web pages as Documents using Spider AI.
Must have the Python package spider-client installed and a Spider API key.
See https://spider.cloud for more.
Load Microsoft Excel files using Unstructured.
Like other Unstructured loaders, UnstructuredExcelLoader can be used in both "single" and "elements" mode. If you use the loader in "elements" mode, e
Load blobs from cloud URL or file:.
Example:
.. code-block:: python
loader = CloudBlobLoader("s3://mybucket/id")
for blob in loader.yield_blobs():
print(blob)
Load YouTube urls as audio file(s).
Load blobs in the local file system.
Example:
.. code-block:: python
from langchain_community.document_loaders.blob_loaders import FileSystemBlobLoader
loader = FileSystemBlobLoader("/path/
Parse the Microsoft Word documents from a blob.
Parse a blob from a PDF using pypdf library.
This class provides methods to parse a blob from a PDF document, supporting various configurations such as handling password-protected PDFs, extra
Parse a blob from a PDF using pdfminer.six library.
This class provides methods to parse a blob from a PDF document, supporting various configurations such as handling password-protected PDFs
Parse a blob from a PDF using PyMuPDF library.
This class provides methods to parse a blob from a PDF document, supporting various configurations such as handling password-protected PDFs, ext
Parse a blob from a PDF using PyPDFium2 library.
This class provides methods to parse a blob from a PDF document, supporting various configurations such as handling password-protected PDFs, e
Parse PDF with PDFPlumber.
Send PDF files to Amazon Textract and parse them.
For parsing multi-page PDFs, they have to reside on S3.
The AmazonTextractPDFLoader calls the [Amazon Textract Service](https://aws.amazon.com/t
Loads a PDF with Azure Document Intelligence (formerly Form Recognizer) and chunks at character level.
Transcribe and parse audio files using Azure OpenAI Whisper.
This parser integrates with the Azure OpenAI Whisper model to transcribe audio files. It differs from the standard OpenAI Whisper parser,
Transcribe and parse audio files.
Audio transcription is with OpenAI Whisper model.
Transcribe and parse audio files with OpenAI Whisper model.
Audio transcription with OpenAI Whisper model locally from transformers.
Transcribe and parse audio files. Audio transcription is with OpenAI Whisper model.
Transcribe and parse audio files with faster-whisper.
faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is up to 4 times faster than openai/whisper for the same
A wrapper class that adapts a document loader to function as a parser.
This class is a work-around that adapts a document loader to function as a parser. It is recommended to use a proper parser, if
Parser that uses mime-types to parse a blob.
This parser is useful for simple pipelines where the mime-type is sufficient to determine how to parse a blob.
To use, configure handlers based on mime
Dataclass to store Document AI parsing results.
Loads a PDF with Azure Document Intelligence (formerly Forms Recognizer).
Parser for text blobs.
Parser for vsdx files.
Abstract base class for parsing image blobs into text.
Parser for extracting text from images using the RapidOCR library.
Parse for extracting text from images using the Tesseract OCR library.
Parser for analyzing images using a language model (LLM).
Exception raised when the Grobid server is unavailable.
Load article PDF files using Grobid.
Code segmenter for Go.
Code segmenter for PHP.
Parse using the respective programming language syntax.
Each top-level function and class in the code is loaded into separate documents. Furthermore, an extra document is generated, containing the re
Code segmenter for C.
Code segmenter for Lua.
Code segmenter for Scala.
Code segmenter for Ruby.
Code segmenter for TypeScript.
Code segmenter for SQL. This class uses Tree-sitter to segment SQL code into its constituent statements (e.g., SELECT, CREATE TABLE). It also provides functionality to extract these statements and sim
Code segmenter for Python.
Code segmenter for C#.
Code segmenter for COBOL.
Abstract class for the code segmenter.
Code segmenter for Java.
Code segmenter for Elixir.
Code segmenter for JavaScript.
Abstract class for CodeSegmenters that use the tree-sitter library.
Code segmenter for Perl.
Code segmenter for Kotlin.
Code segmenter for Rust.
Code segmenter for C++.
Parse HTML files using Beautiful Soup.
Document compressor that uses Volcengine Rerank API.
Document compressor using Flashrank interface.
Document compressor that uses Jina Rerank API.
Request for reranking.
OpenVINO rerank models.
Document compressor using Flashrank interface.
Compress using LLMLingua Project.
https://github.com/microsoft/LLMLingua
Document compressor that uses Infinity Rerank API.
Document compressor that uses DashScope Rerank API.
Transform HTML content by extracting specific tags and removing unwanted ones.
Nuclia Text Transformer.
The Nuclia Understanding API splits into paragraphs and sentences, identifies entities, provides a summary of the text and generates embeddings for all sentences.
Replace occurrences of a particular search pattern with a replacement string
Reorder long context.
Lost in the middle: Performance degrades when models must access relevant information in the middle of long contexts. See: https://arxiv.org/abs//2307.03172
Extract metadata tags from document contents using OpenAI functions.
Example: .. code-block:: python
from langchain_openai import ChatOpenAI
from langchai
Extract properties from text documents using doctran.
Translate text documents using doctran.
Converts HTML documents to Markdown format with customizable options for handling links, images, other tags and heading styles using the markdownify library.
Extract QA from text documents using doctran.
Filter that drops redundant documents by comparing their embeddings.
Perform K-means clustering on document vectors. Returns an arbitrary number of documents closest to center.
Load telegram conversations to LangChain chat messages.
To export, use the Telegram Desktop app from https://desktop.telegram.org/, select a conversation, click the three dots in the top right corn
Load chat sessions from a list of LangSmith "llm" runs.
Load chat sessions from a LangSmith dataset with the "chat" data type.
Load chat sessions from the iMessage chat.db SQLite file.
It only works on macOS when you have iMessage enabled and have the chat.db file.
The chat.db file is likely located at ~/Library/Messages/
Load Slack conversations from a dump zip file.
Load WhatsApp conversations from a dump zip file or directory.
Load Facebook Messenger chat data from a single file.
Load Facebook Messenger chat data from a folder.
Load chat sessions from Gmail.
Inherits from
BaseChatLoader.
Loads sent messages and their preceding emails to create chat training examples.
This lo
Load from oracle adb
Autonomous Database connection can be made by either connection_string or tns name. wallet_location and wallet_password are required for TLS connection. Each document will repres
Read documents using OracleDocLoader Args: conn: Oracle Connection, params: Loader parameters.
Splitting text using Oracle chunker.
Load from Azure Blob Storage container.
Load from the Google Cloud Platform BigQuery.
Each document represents one row of the result. The page_content_columns
are written into the page_content of the document. The metadata_columns
Load from GCS file.
Load Google Docs from Google Drive.
Load from GCS directory.
Loader for Google Cloud Speech-to-Text audio transcripts.
It uses the Google Cloud Speech-to-Text API to transcribe audio files and loads the transcribed text into one or more Documents, depending on
Load from Azure Blob Storage files.
Load datasets from Apify web scraping, crawling, and data extraction platform.
For details, see https://docs.apify.com/platform/integrations/langchain
Load files using Unstructured.
The file loader uses the unstructured partition function and will automatically detect the file type. You can run the loader in different modes: "single", "elements",
Load files using Unstructured API.
By default, the loader makes a call to the hosted Unstructured API. If you are running the unstructured API locally, you can change the API rule by passing in the
Load file-like objects opened in read mode using Unstructured.
The file loader uses the unstructured partition function and will automatically detect the file type. You can run the loader in differ
Send file-like objects with unstructured-client sdk to the Unstructured API.
By default, the loader makes a call to the hosted Unstructured API. If you are running the unstructured API locally, you
Load from Docugami.
To use, you should have the dgml-utils python package installed.
Google Cloud Document AI parser.
For a detailed explanation of Document AI, refer to the product documentation. https://cloud.google.com/document-ai/docs/overview
Translate text documents using Google Cloud Translation.