Releases · Unstructured-IO/unstructured

06 Feb 06:12

cragwolfe

0.16.20

b10379c

0.16.20 Latest

Latest

0.16.20

Enhancements

Features

Fixes

Fix a security issue where rst and org files could read files in the local filesystem. Certain filetypes could 'include' or 'import' local files into their content, allowing partitioning of arbitrary files from the local filesystem. Partitioning of these files is now sandboxed.

Assets 2

05 Feb 17:21

plutasnyy

0.16.19

5852260

0.16.19

Enhancements

Features

Fixes

Fix a bug where table extraction is skipped when it shouldn't. Pages with just one table as its content or starts with a table misses table extraction. The routing logic is now fixed.
Correct deprecated ruff invocation in make tidy. This will future-proof it or avoid surprises if someone happens to upgrade Ruff.
Remove upper bound constraint on python version in setup.py. Python3.13 is not yet officially supported, but allow users to try.
Fixes removing HTML elements from the inside of table cells in html partition v=2.0. The HTML partitioner now correctly preserves HTML elements from the inside of table cells.

Assets 2

29 Jan 12:52

badGarnet

0.16.17

55debaf

0.16.17

Enhancements

Refactoring the VoyageAI integration to use voyageai package directly, allowing extra features.

Features

Fixes

Fix a bug where build_layout_elements_from_cor_regions incorrectly joins texts in wrong order.

Full Changelog: 0.16.16...0.16.17

Assets 2

27 Jan 23:30

christinestraub

0.16.16

a447b81

0.16.16

Enhancements

Features

Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.

Drop usage of ndjson dependency

Assets 2

23 Jan 04:51

tbs17

0.16.15

8d0b68a

0.16.15

Update unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libs
Update pdfminer-six to 20240706

Assets 2

20 Jan 13:00

plutasnyy

0.16.14

efd9f64

0.16.14

Enhancements

Features

Fixes

Fix an issue with multiple values for infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure already

Assets 2

13 Jan 15:40

plutasnyy

0.16.13

38eb661

0.16.13

Enhancements

Add character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.

Features

Fixes

Fix NLTK Download to use nltk assets in docker image
removed the ability to automatically download nltk package if missing

Assets 2

05 Jan 22:06

cragwolfe

0.16.12

1a94d95

0.16.12

Enhancements

Prepare auto-partitioning for pluggable partitioners. Move toward a uniform partitioner call signature so a custom or override partitioner can be registered without code changes.
Add NDJSON file type support.

Features

Fixes

Base image has been updated.
Upgrade ruff to latest. Previously the ruff version was pinned to <0.5. Remove that pin and fix the handful of lint items that resulted.
CSV with asserted XLS content-type is correctly identified as CSV. Resolves a bug where a CSV file with an asserted content-type of application/vnd.ms-excel was incorrectly identified as an XLS file.
Improve element-type mapping for Chinese text. Fixes bug where Chinese text would produce large numbers of false-positive Title elements.
Improve element-type mapping for HTML. Fixes bug where certain non-title elements were classified as Title.

Assets 2

10 Dec 00:51

scanny

0.16.11

b981d71

0.16.11

Enhancements

Enhance quote standardization tests with additional Unicode scenarios
Relax table segregation rule in chunking. Previously a Table element was always segregated into its own pre-chunk such that the Table appeared alone in a chunk or was split into multiple TableChunk elements, but never combined with Text-subtype elements. Allow table elements to be combined with other elements in the same chunk when space allows.
Compute chunk length based solely on element.text. Previously .metadata.text_as_html was also considered and since it is always longer that the text (due to HTML tag overhead) it was the effective length criterion. Remove text-as-html from the length calculation such that text-length is the sole criterion for sizing a chunk.

Features

Fixes

Fix ipv4 regex to correctly include up to three digit octets.

Assets 2

07 Dec 18:13

tbs17

0.16.10

59e6cff

0.16.10

Enhancements

Features

Fixes

Fix original file doctype detection from cct converted file paths for metrics calculation.

Assets 2

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

0.16.20

Enhancements

Features

Fixes

Enhancements

Features

Fixes

0.16.17

Enhancements

Features

Fixes

0.16.16

Enhancements

Features

Fixes

Enhancements

Features

Fixes

Enhancements

Features

Fixes

0.16.12

Enhancements

Features

Fixes

Enhancements

Features

Fixes

0.16.10

Enhancements

Features

Fixes

Releases: Unstructured-IO/unstructured

0.16.20

0.16.20

Enhancements

Features

Fixes

0.16.19

Enhancements

Features

Fixes

0.16.17

0.16.17

Enhancements

Features

Fixes

0.16.16

0.16.16

Enhancements

Features

Fixes

0.16.15

0.16.14

Enhancements

Features

Fixes

0.16.13

Enhancements

Features

Fixes

0.16.12

0.16.12

Enhancements

Features

Fixes

0.16.11

Enhancements

Features

Fixes

0.16.10

0.16.10

Enhancements

Features

Fixes