Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 21 additions & 21 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -262,13 +262,13 @@ All notable changes to this project will be documented in this file.
```py
from presidio_analyzer import AnalyzerEngine
from presidio_analyzer.predefined_recognizers import AuAbnRecognizer

# Initialize an analyzer engine with the recognizer registry
analyzer = AnalyzerEngine()

# Create an instance of the AuAbnRecognizer
au_abn_recognizer = AuAbnRecognizer()

# Add the recognizer to the registry
analyzer.registry.add_recognizer(au_abn_recognizer)
```
Expand Down Expand Up @@ -310,14 +310,14 @@ All notable changes to this project will be documented in this file.
- Drop dependency of spacy_stanza package, and add supporting code to stanza_nlp_engine, to support recent stanza versions
- Add parameters to allow users to define the number of processes and batch size when running BatchAnalyzerEngine.
- Fix InPassportRecognizer regex recognizer

### Anonymizer
- Changed: Deprecate `MD5` hash type option, defaulting into `sha256`.
- Replace crypto package dependency from pycryptodom to cryptography
- Replace crypto package dependency from pycryptodom to cryptography
- Remove azure-core dependency from anonymizer

### Image Redactor
- Changed: Updated the return type annotation of `ocr_bboxes` in `verify_dicom_instance()` from `dict` to `list`.
- Changed: Updated the return type annotation of `ocr_bboxes` in `verify_dicom_instance()` from `dict` to `list`.

### Presidio Structured

Expand Down Expand Up @@ -383,15 +383,15 @@ All notable changes to this project will be documented in this file.
* Add a link to HashiCorp vault operator resource ([#1468](https://github.com/data-privacy-stack/presidio/pull/1468)) (Thanks [Akshay Karle](https://github.com/akshaykarle))
### Changed
#### Analyzer
* Updates to the transformers conf docs and yaml file ([#1467](https://github.com/data-privacy-stack/presidio/pull/1467))
* Updates to the transformers conf docs and yaml file ([#1467](https://github.com/data-privacy-stack/presidio/pull/1467))

#### Docs
* docs: clarify the docs on deploying presidio to k8s ([#1453](https://github.com/data-privacy-stack/presidio/pull/1453)) (Thanks [Roel Fauconnier](https://github.com/roel4ez))


## [2.2.355] - July 9th 2024

> Note: A new YAML based mechanism has been added to support no-code customization and creation of recognizers.
> Note: A new YAML based mechanism has been added to support no-code customization and creation of recognizers.
The default recognizers are now automatically loaded from file.


Expand Down Expand Up @@ -422,7 +422,7 @@ The default recognizers are now automatically loaded from file.
* Add Ruff linter + Apply Ruff fix (#1379)
* Auto-formatting, fix D rules (#1381)
* Fix N818, E721 (#1382)
* Migrate Python Packaging to pyproject.toml (#1383)
* Migrate Python Packaging to pyproject.toml (#1383)
* From Pipenv to Poetry (#1391)
* Fix ports in docs (#1408)

Expand All @@ -445,7 +445,7 @@ The default recognizers are now automatically loaded from file.

#### Image redactor
* Added tesseract to installation (#1312)

#### Structured
* Analysis builder improvements (#1295) (Thanks @ebotiab)
* Implement user-defined entity selection strategies in Presidio Structured (#1319) (Thanks @miltonsim)
Expand Down Expand Up @@ -541,10 +541,10 @@ The default recognizers are now automatically loaded from file.
### Added
#### Analyzer
* New Predefined Recognizer: IN_PAN (#1100)

#### Anonymizer
* Anonymizer - Pass bytes key to Encrypt / Decrypt (#1147)

#### Image redactor
* DICOM redactor improvement: Enabling more photometric interpretations (#1103)
* DICOM redactor improvement: Adding exceptions for when DICOM file does not have pixel data (#1104)
Expand All @@ -563,27 +563,27 @@ The default recognizers are now automatically loaded from file.
* Updating verification engines and enable plotting of custom bboxes (#1164)
* Added image processing class to preprocess the image before running OCR (#1166)
* Added support for Microsoft's document intelligence OCR

### Changed
#### Analyzer
* Refactored the `NlpEngine` and Ner recognizers (`SpacyRecognizer`, `TransformersRecognizer`, `StanzaRecognizer`) to allow simpler integration of huggingface and transformers models (#1159). This includes:
- Changes in how NER results flow through Presidio (see docs)
- NER/model definition is now defined using a conf file or a `NerModelConfiguration` object.
- Integrated `spacy-huggingface-pipelines` for a more robust integration of huggingface models.
- Integrated `spacy-huggingface-pipelines` for a more robust integration of huggingface models.
* As a result, `SpacyRecognizer` logic has changed, please see #1159. Some fields within the class are now deprecated.
* Updated type checks (#1175)
* Enabled regex flags manipulation (#1193)

#### Anonymizer
* Initial logic check for merging 2 entities (#1092)
* Fix Sphinx warning in OperatorConfig (#1143)
* Fix type mismatch in check_label_groups parameter in spacy_recognizer (#1130)
* anonymize_list return type hint fix (#1178)
*
*
#### General
* We no longer use Pipenv.lock. Locking happens as part of the CI. (#1152)
* Changed the ACR instance (#1089)
* Updated to Cred Scan V3 (#1154)
* Updated to Cred Scan V3 (#1154)

## [2.2.33] - June 1st 2023
### Added
Expand Down Expand Up @@ -723,7 +723,7 @@ Bug fix in context support

* Added a URL recognizer
* Added a new capability for creating new logic for context detection. See [ContextAwareEnhancer](https://github.com/data-privacy-stack/presidio/blob/main/presidio-analyzer/presidio_analyzer/context_aware_enhancers/context_aware_enhancer.py) and [LemmaContextAwareEnhancer](https://github.com/data-privacy-stack/presidio/blob/main/presidio-analyzer/presidio_analyzer/context_aware_enhancers/lemma_context_aware_enhancer.py). Documentation would be added on a future release.
Furthermore, it is now possible to pass context words thruogh the `analyze` method (or via API) and those would be taken into account for context enhancement.
Furthermore, it is now possible to pass context words thruogh the `analyze` method (or via API) and those would be taken into account for context enhancement.



Expand Down Expand Up @@ -808,7 +808,7 @@ Upgrade Analyzer spacy version to 3.0.5
- In OperatorConfig: anonymizer_name -> operator_name
2. Response entity AnonymizerResult renamed to EngineResult
- In EngineResult: List[AnonymizedEntity] -> List[OperatorResult]
- In OperatorResult:
- In OperatorResult:
- anonymizer -> operator
- anonymized_text -> text

Expand All @@ -817,7 +817,7 @@ Upgrade Analyzer spacy version to 3.0.5
2. Response entity anonymizer_text renamed to text.

#### Deanonymize:
New endpoint for deanonymizing encrypted entities by the anonymizer.
New endpoint for deanonymizing encrypted entities by the anonymizer.


[unreleased]: https://github.com/data-privacy-stack/presidio/compare/2.2.363...HEAD
Expand Down
2 changes: 1 addition & 1 deletion SUPPORT.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
# Support
This open source project does not have any SLA or official support.

The best way to get support is to open an issue or start a discussion.
The best way to get support is to open an issue or start a discussion.
Before doing so, consider searching on the [Github repo](https://github.com/search?q=repo%3Adata-privacy-stack%2Fpresidio%20&type=code) or the [docs website](https://data-privacy-stack.github.io/presidio/).
2 changes: 1 addition & 1 deletion docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ services:
interval: 10s
timeout: 10s
retries: 30 # ~5 minutes total
start_period: 60s
start_period: 60s
presidio-anonymizer:
image: ${REGISTRY_NAME}/${IMAGE_PREFIX}presidio-anonymizer${TAG}
build:
Expand Down
8 changes: 4 additions & 4 deletions docs/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,9 +44,9 @@ with at least one NLP engine (`spaCy`, `transformers` or `stanza`):
```

!!! note "Note"

When using a transformers NLP engine, Presidio would still use spaCy for other capabilities,
therefore a small spaCy model (such as en_core_web_sm) is required.
therefore a small spaCy model (such as en_core_web_sm) is required.
Transformers models would be loaded lazily. To pre-load them, see: [Downloading a pre-trained model](./analyzer/nlp_engines/transformers.md#downloading-a-pre-trained-model)

=== "Stanza"
Expand All @@ -58,7 +58,7 @@ with at least one NLP engine (`spaCy`, `transformers` or `stanza`):


!!! note "Note"

Stanza models would be loaded lazily. To pre-load them, see: [Downloading a pre-trained model](./analyzer/nlp_engines/spacy_stanza.md#download-the-pre-trained-model).

### GPU acceleration (optional)
Expand All @@ -78,7 +78,7 @@ For PII redaction in images

```sh
pip install presidio_image_redactor

# Presidio image redactor uses the presidio-analyzer
# which requires a spaCy language model:
python -m spacy download en_core_web_lg
Expand Down
6 changes: 3 additions & 3 deletions docs/presidio_V2.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,12 @@ The main changes introduced in **V2** are:

3. Improved documentation, sample code, and build workflows.

4. Format-Preserving Encryption replaced with Advanced Encryption Standard (AES)
4. Format-Preserving Encryption replaced with Advanced Encryption Standard (AES)

## V1 Availability

Version V1 (legacy) is still available for download. To continue using the previous version:
- For docker containers, use tag=v1
- For docker containers, use tag=v1
- For python packages, download version < 2 (e.g. pip install presidio-analyzer==0.95)

!!! note "Note"
Expand All @@ -33,7 +33,7 @@ The move from gRPC to HTTP-based APIs included changes to the API requests.
1. Changed payload format – moving from structured objects to JSON.

2. Removed templates from the API, including flattening the JSON structure.

3. Using snake_case instead of camelCase .

Below is a detailed outline of all changes made to the Analyzer and Anonymizer.
Expand Down
8 changes: 4 additions & 4 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,19 +8,19 @@ repo_url: https://github.com/data-privacy-stack/presidio/
edit_uri: ""

nav:
- Presidio:
- Presidio:
- Home: index.md
- Installation: installation.md
- FAQ: faq.md
- Quick start:
- Quick start:
- Home: getting_started.md
- Text: getting_started/getting_started_text.md
- Images: getting_started/getting_started_images.md
- Semi/Structured data: getting_started/getting_started_structured.md
- Learn Presidio:
- Home: learn_presidio/index.md
- Concepts: learn_presidio/concepts.md
- Tutorial:
- Tutorial:
- Home: tutorial/index.md
- Getting started: tutorial/00_getting_started.md
- Deny-list recognizers: tutorial/01_deny_list.md
Expand Down Expand Up @@ -241,7 +241,7 @@ markdown_extensions:
- pymdownx.superfences
- pymdownx.pathconverter
- pymdownx.tabbed:
alternate_style: true
alternate_style: true
- pymdownx.superfences:
custom_fences:
- name: mermaid
Expand Down