You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
There is no in-place upgrade from FSCrawler 2.9. Install 3.0, recreate jobs with --setup, and reindex. See {ref}upgrade-from-2.9.
Breaking changes
If you want to exclude a specific folder, you need to use a wildcard character at the end of the folder name.
For example, to exclude the folder /tmp/foo, you need to use /tmp/foo/*. Thanks to dadoonet.
The way we run docker images has changed. We don't need anymore to specify the fscrawler binary.
So running docker run -it -v ~/.fscrawler:/root/.fscrawler -v /documents:/tmp/es:ro dadoonet/fscrawler job_name is
enough. Thanks to dadoonet.
FSCrawler does not display anymore the list of existing jobs when no job name is provided.
You need to use the --list option to list the jobs. Thanks to dadoonet.
When launching for the first time FSCrawler with a job name, FSCrawler does not create anymore the job
configuration folder with default settings. You need to use the --setup option to create the job settings.
Thanks to dadoonet.
We don't support anymore the elasticsearch.nodes.url setting. You need to use elasticsearch.urls
instead. Thanks to dadoonet.
The _upload REST endpoint has been removed. Please now use the _document endpoint. Thanks to dadoonet.
The Apache Tika 4 upgrade changes several meta.raw.* metadata key names compared to FSCrawler 2.9.
If you search, aggregate or map these fields, you will need to update your queries and index templates
accordingly:
Image and EXIF metadata (JPEG, PNG, TIFF, …) keys are now namespaced under img:. For example Number of Tables becomes img:Number of Tables and Exif IFD0:Orientation becomes img:Exif IFD0:Orientation.
ICC profile keys use the lowercase icc: prefix instead of ICC:.
PDF access-permission:* and pdf:* keys now use hyphens instead of underscores or camelCase.
For example access_permission:assemble_document becomes access-permission:assemble-document
and pdf:PDFVersion becomes pdf:pdf-version.
The resource name key resourceName is renamed to tk:resource-name.
Tika internal keys now use the tk: prefix instead of X-TIKA:. New keys include tk:content-type-magic-detected, tk:parsed-by-full-set and, for text documents, the encoding
detection keys tk:detected-encoding, tk:encoding-detection-trace and tk:encoding-detector.
Thanks to dadoonet.
The default fs.ocr.pdf_strategy is now auto instead of ocr_and_text. With auto, OCR is skipped on
PDF pages that already contain more than 10 characters of text. If you relied on the previous behaviour and want
OCR on every PDF page, explicitly set fs.ocr.pdf_strategy to ocr_and_text. Thanks to dadoonet.
New jobs created with --setup set fs.hash_algorithm to SHA-256 in the example settings. Existing jobs that
omit the setting keep MD5 so document _ids stay unchanged. Changing the algorithm later requires a full
reindex. See document ids. Closes #2425. Thanks to dadoonet.
New
Default fs.excludes now also skips macOS Finder metadata files (.DS_Store), via the
case-insensitive pattern */.ds_store, in addition to */~*. See includes_excludes.
Thanks to dadoonet.
The crawler system has been unified using a plugin architecture. You can now specify the crawler provider using fs.provider instead of server.protocol. Available providers are local (default), ftp, and ssh.
See crawler provider. Thanks to dadoonet.
FSCrawler does not need to wait until the next planned scan to scan again the filesystem. You can just set the next_check field to null in the ~/.fscrawler/{job_name}/_checkpoint.json file and FSCrawler will start
a new scan immediately.
Job settings can be defined by env variables and system properties and you can also split the configuration of
jobs using multiple files in the ~/.fscrawler/job/_settings directory. Also note that the system properties
need to be set in the FS_JAVA_OPTS environment variable.
Add support for automatic semantic search when using a 8.17+ version with a trial or enterprise
license. See semantic_search. Warning: this might slow down the ingestion process. Thanks to dadoonet.
Add support for Elastic cloud serverless. Thanks to dadoonet.
Using the REST API _document, you can now fetch a document from the local dir, from an http website
or from an S3 bucket. See rest service. Thanks to dadoonet.
You can now remove a document in Elasticsearch using FSCrawler _document endpoint. See rest service. Thanks to dadoonet.
Implement our own HTTP Client for Elasticsearch. Thanks to dadoonet.
FSCrawler now ships with Apache Tika's tika-vlm module: OCR can be delegated to a Vision Language
Model through an OpenAI-compatible endpoint (vLLM, Ollama, Azure OpenAI…), Anthropic Claude or Google
Gemini, configured via a custom Tika configuration file. See vlm ocr. The default FSCrawler
parser chain is unchanged (Tesseract when available). Thanks to dadoonet.
Add option to set path to custom tika config file. See local fs settings. Thanks to iadcode for the original
XML implementation and to betofilippi for the switch to JSON.
If your JSON configuration uses default-parser, exclude the VLM parser components unless you
explicitly enable one — see local fs settings and vlm ocr. Note: since the Apache Tika 4 upgrade, the configuration file must be a Tika JSON configuration —
the XML-based configuration file mechanism was removed upstream. Existing XML configurations need to be
converted. See local fs settings.
Support for Index Templates. See mappings. Thanks to dadoonet.
Support for Aliases. You can now index to an alias. Thanks to dadoonet.
Support for Access Token and Api Keys instead of Basic Authentication. See credentials. Thanks to dadoonet.
Allow loading external jars. This adds a new external directory from where jars can be loaded
to the FSCrawler JVM. For example, you could provide your own Custom Tika Parser code. See layout. Thanks to dadoonet.
Add temporal information in folder index. Thanks to bdauvissat
Add support for external metadata files while crawling, defaults to .meta.yml. See tags Thanks to dadoonet.
Add support for static external metadata for all documents. See tags Thanks to dadoonet.
The job name is not mandatory anymore and it will be fscrawler by default. Thanks to dadoonet.
FSCrawler also supports Elasticsearch 9. Thanks to dadoonet.
Add support for ACL metadata extraction for NTFS filesystems, including principals, permissions, and flags. Thanks to alexbluesteele.
Add support for pause/resume functionality with checkpoint persistence. The crawler can now be paused and resumed
without losing progress. It also automatically recovers from network errors with exponential backoff retry.
See rest service. Thanks to dadoonet.
HTTP retry backoff is configurable via elasticsearch.retry_max_duration, elasticsearch.retry_initial_delay and elasticsearch.retry_max_delay
(defaults 5m / 500ms / 30s). The same budget applies to 5xx, 429, and a
cold-start 404 on GET /. See http retry settings. Thanks to dadoonet.
FSCrawler can create a default Kibana dashboard on job startup via the Kibana Dashboards API (Kibana 9.5+).
See kibana settings. Closes #2477. Thanks to dadoonet.
Fix
Apple Keynote (.key) files are now supported for content extraction and indexing. Closes #782.
Closed open file streams after use. Thanks to alexbluesteele.
fs.ocr.enabled was always false. Thanks to ywjung.
Do not hide YAML parsing errors. Thanks to dadoonet.
Fix duration parsing for the day unit d. Thanks to dadoonet.
Image raw metadata extraction was not working. Thanks to dadoonet.
Fix issue when using crawling over SSH when the directory ends with a space. Thanks to dadoonet.
On Windows, files and directories to be removed were not properly detected. Thanks to newschapmj1.
Bulk _bulk HTTP calls now retry on 429/5xx and no longer treat a failed bulk as success.
Exhausted retries mark the crawl checkpoint as ERROR and REST uploads return ok: false.
Thanks to dadoonet.
Default Log4J config sets org.apache.pdfbox to error to avoid flooding logs with No Unicode mapping for CID+… warnings from subset fonts. Thanks to dadoonet.
Default Log4J config sets org.apache.fontbox to error to avoid flooding logs with No PostScript name data is provided for the font … warnings. Thanks to dadoonet.
Failed bulk actions are detailed in logs/bulk-failures.log (reason-prefixed lines; truncated
payloads at TRACE). Console / fscrawler.log point to that file. Thanks to dadoonet.
Deprecated
The server.protocol setting is deprecated. Use fs.provider instead. Thanks to dadoonet.
Support for Basic Authentication is deprecated. You should use API keys instead. Thanks to dadoonet.
Updated
Files are now sorted by date with a reverse order. So the most recent files should be indexed first. Thanks to dadoonet.
Add full support for Elasticsearch 9.5.2, 8.19.5, 7.17.29. Thanks to dadoonet.
The default alias name is now the job name and not forced to fscrawler anymore. Thanks to dadoonet.
The default REST endpoint is now running at / instead of /fscrawler/. Thanks to dadoonet.
Upgrade to Jackson 3.x. Closes #2419. Thanks to dadoonet.
As a consequence of the Jackson 3 upgrade, the JSON and YAML documents produced by FSCrawler now serialize their
fields in alphabetical order. This affects the indexed documents, the generated _settings.yaml and _checkpoint.json files, and the REST API responses. This is purely cosmetic (field order is not significant in
JSON), but users who version their configuration files may notice a one-time reordering. Thanks to dadoonet.
Removed
Remove the specific distributions depending on Elastic version. Thanks to dadoonet.
Support for Elasticsearch 6.x is removed. Thanks to dadoonet.
Thanks to @dadoonet, @ywjung, @iadcode, @bdauvissat, @alexbluesteele, @betofilippi, @newschapmj1
and all the contributors for this release!
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
The FSCrawler team is pleased to announce the FSCrawler 3.0 release!
FSCrawler is a crawler for Elasticsearch that helps to index binary documents such as PDF, MS Office, and more.
Usage
Download FSCrawler 3.0:
wget https://repo1.maven.org/maven2/fr/pilato/elasticsearch/crawler/fscrawler-distribution/3.0/fscrawler-distribution-3.0.zip unzip fscrawler-distribution-3.0.zip cd fscrawler-distribution-3.0On first run, create the default job configuration:
Create a directory such as
/tmp/es, add files to index, then start FSCrawler:Or with Docker:
docker run -it --rm \ -v ~/.fscrawler:/root/.fscrawler \ -v ~/tmp:/tmp/es:ro \ dadoonet/fscrawler:3.0On first run with Docker, add
--setupto create the configuration.More details in the documentation.
Version 3.0
There is no in-place upgrade from FSCrawler 2.9. Install 3.0, recreate jobs with
--setup, and reindex. See {ref}upgrade-from-2.9.Breaking changes
For example, to exclude the folder
/tmp/foo, you need to use/tmp/foo/*. Thanks to dadoonet.So running
docker run -it -v ~/.fscrawler:/root/.fscrawler -v /documents:/tmp/es:ro dadoonet/fscrawler job_nameisenough. Thanks to dadoonet.
You need to use the
--listoption to list the jobs. Thanks to dadoonet.configuration folder with default settings. You need to use the
--setupoption to create the job settings.Thanks to dadoonet.
elasticsearch.nodes.urlsetting. You need to useelasticsearch.urlsinstead. Thanks to dadoonet.
_uploadREST endpoint has been removed. Please now use the_documentendpoint. Thanks to dadoonet.meta.raw.*metadata key names compared to FSCrawler 2.9.If you search, aggregate or map these fields, you will need to update your queries and index templates
accordingly:
img:. For exampleNumber of Tablesbecomesimg:Number of TablesandExif IFD0:Orientationbecomesimg:Exif IFD0:Orientation.icc:prefix instead ofICC:.access-permission:*andpdf:*keys now use hyphens instead of underscores or camelCase.For example
access_permission:assemble_documentbecomesaccess-permission:assemble-documentand
pdf:PDFVersionbecomespdf:pdf-version.resourceNameis renamed totk:resource-name.tk:prefix instead ofX-TIKA:. New keys includetk:content-type-magic-detected,tk:parsed-by-full-setand, for text documents, the encodingdetection keys
tk:detected-encoding,tk:encoding-detection-traceandtk:encoding-detector.Thanks to dadoonet.
fs.ocr.pdf_strategyis nowautoinstead ofocr_and_text. Withauto, OCR is skipped onPDF pages that already contain more than 10 characters of text. If you relied on the previous behaviour and want
OCR on every PDF page, explicitly set
fs.ocr.pdf_strategytoocr_and_text. Thanks to dadoonet.--setupsetfs.hash_algorithmtoSHA-256in the example settings. Existing jobs thatomit the setting keep
MD5so document_ids stay unchanged. Changing the algorithm later requires a fullreindex. See document ids. Closes #2425. Thanks to dadoonet.
New
fs.excludesnow also skips macOS Finder metadata files (.DS_Store), via thecase-insensitive pattern
*/.ds_store, in addition to*/~*. See includes_excludes.Thanks to dadoonet.
fs.providerinstead ofserver.protocol. Available providers arelocal(default),ftp, andssh.See crawler provider. Thanks to dadoonet.
next_checkfield tonullin the~/.fscrawler/{job_name}/_checkpoint.jsonfile and FSCrawler will starta new scan immediately.
jobs using multiple files in the
~/.fscrawler/job/_settingsdirectory. Also note that the system propertiesneed to be set in the
FS_JAVA_OPTSenvironment variable.license. See semantic_search. Warning: this might slow down the ingestion process. Thanks to dadoonet.
_document, you can now fetch a document from the local dir, from an http websiteor from an S3 bucket. See rest service. Thanks to dadoonet.
_documentendpoint. See rest service. Thanks to dadoonet.tika-vlmmodule: OCR can be delegated to a Vision LanguageModel through an OpenAI-compatible endpoint (vLLM, Ollama, Azure OpenAI…), Anthropic Claude or Google
Gemini, configured via a custom Tika configuration file. See vlm ocr. The default FSCrawler
parser chain is unchanged (Tesseract when available). Thanks to dadoonet.
XML implementation and to betofilippi for the switch to JSON.
If your JSON configuration uses
default-parser, exclude the VLM parser components unless youexplicitly enable one — see local fs settings and vlm ocr.
Note: since the Apache Tika 4 upgrade, the configuration file must be a Tika JSON configuration —
the XML-based configuration file mechanism was removed upstream. Existing XML configurations need to be
converted. See local fs settings.
externaldirectory from where jars can be loadedto the FSCrawler JVM. For example, you could provide your own Custom Tika Parser code. See layout. Thanks to dadoonet.
.meta.yml. See tags Thanks to dadoonet.fscrawlerby default. Thanks to dadoonet.without losing progress. It also automatically recovers from network errors with exponential backoff retry.
See rest service. Thanks to dadoonet.
elasticsearch.retry_max_duration,elasticsearch.retry_initial_delayandelasticsearch.retry_max_delay(defaults
5m/500ms/30s). The same budget applies to5xx,429, and acold-start
404onGET /. See http retry settings. Thanks to dadoonet.See kibana settings. Closes #2477. Thanks to dadoonet.
Fix
.key) files are now supported for content extraction and indexing. Closes #782.fs.ocr.enabledwas always false. Thanks to ywjung.d. Thanks to dadoonet._bulkHTTP calls now retry on429/5xxand no longer treat a failed bulk as success.Exhausted retries mark the crawl checkpoint as
ERRORand REST uploads returnok: false.Thanks to dadoonet.
org.apache.pdfboxtoerrorto avoid flooding logs withNo Unicode mapping for CID+…warnings from subset fonts. Thanks to dadoonet.org.apache.fontboxtoerrorto avoid flooding logs withNo PostScript name data is provided for the font …warnings. Thanks to dadoonet.logs/bulk-failures.log(reason-prefixed lines; truncatedpayloads at TRACE). Console /
fscrawler.logpoint to that file. Thanks to dadoonet.Deprecated
server.protocolsetting is deprecated. Usefs.providerinstead. Thanks to dadoonet.Updated
fscrawleranymore. Thanks to dadoonet./instead of/fscrawler/. Thanks to dadoonet.fields in alphabetical order. This affects the indexed documents, the generated
_settings.yamland_checkpoint.jsonfiles, and the REST API responses. This is purely cosmetic (field order is not significant inJSON), but users who version their configuration files may notice a one-time reordering. Thanks to dadoonet.
Removed
Thanks to
@dadoonet,@ywjung,@iadcode,@bdauvissat,@alexbluesteele,@betofilippi,@newschapmj1and all the contributors for this release!
What's Changed
fs.ocr.enabledis always false by @ywjung infs.ocr.enabledis always false #1358fileinstead ofFilefor facets by @dadoonet in Usefileinstead ofFilefor facets #1552--debugand--traceby @dadoonet in Deprecate--debugand--trace#1780elasticsearch.byte_sizeby @dadoonet in Supportelasticsearch.byte_size#1852filesizeis not provided by curl by @dadoonet infilesizeis not provided by curl #1871mvn testby @dadoonet in Unit tests not picked up bymvn test#1928build_flavorcheck. by @dadoonet in https://github.com/dadoonet/fscrawler/pull/2075next_checkto_status.jsonby @dadoonet in https://github.com/dadoonet/fscrawler/pull/2106path.*fields by @dadoonet in https://github.com/dadoonet/fscrawler/pull/2198deployphase by @dadoonet in https://github.com/dadoonet/fscrawler/pull/2203elasticsearch.nodes.urlsetting byelasticsearch.urlsby @dadoonet in https://github.com/dadoonet/fscrawler/pull/2206ThisDirHasSpaceAtEndby @dadoonet in https://github.com/dadoonet/fscrawler/pull/2226_iddeduplication in Tips and tricks by @dadoonet in https://github.com/dadoonet/fscrawler/pull/2461New Contributors
fs.ocr.enabledis always false #1358Full Changelog: fscrawler-2.9...fscrawler-3.0
This discussion was created from the release FSCrawler 3.0 🌈.
All reactions