The Filter content GUI page describes the current profile sidebar. This Programming page documents raw values and Java regex execution separately from the user interface.
| Name | Value and meaning |
|---|---|
FILTERS | Contains the REGEXP filter identifier to activate the content filter. Values are stored as a comma-separated list. |
FILTER_PATTERNS | Contains multiple pattern lines separated by literal line breaks. The canonical format of each line is pattern|type|status, with type as text or regexp and status as active or inactive. |
REGEX_MATCH_WORDS | Boolean value, default true. With true, every word touched by a match is removed completely. With false, only the matching character ranges are removed. |
Inactive patterns remain stored but are not applied. Active patterns must contain content. An active regex pattern must compile as a Java pattern; an invalid inactive pattern can remain as a draft.
When a pattern is saved, literal line breaks inside the pattern are encoded as CR. The decoder also understands the legacy token BR. If a pattern itself contains one of these tokens or starts with the reserved prefix, the codec uses the versioned envelope <base64> so that the pattern text remains unambiguous.
<entry key="FILTERS">REGEXP</entry> <entry key="REGEX_MATCH_WORDS">true</entry> <entry key="FILTER_PATTERNS">Version [[CR]][[CR]]|text|active (?s)Start.*?End|regexp|active</entry>
The lines in FILTER_PATTERNS are processed in their stored order. Type text is matched literally; type regexp is compiled as a Java regex.
| Key | Mapping |
|---|---|
pdfc.filter.regex_filter | Adds or removes REGEXP in FILTERS. |
pdfc.filter.regex_filter.filter_patterns | Accesses FILTER_PATTERNS as a list of pattern lines. |
pdfc.filter.regex_filter.match_words | Accesses REGEX_MATCH_WORDS. true corresponds to Whole words; false corresponds to Partial matches. |
Regex patterns are compiled with Java Pattern.compile without a fixed flag mask. Plain-text patterns are quoted with Java Pattern.quote before compilation. Java regular-expression rules therefore apply, including embedded flags.
(?s) enables DOTALL for the following expression and lets the dot match across line breaks.(?m) enables MULTILINE and changes the meaning of ^ and $ for individual lines.(?i) for a case-insensitive search.(?s)Start.*?End (?m)^Article:.*$
Each active pattern is applied with Java Matcher.find() to the complete extracted page text. Matching is not performed separately for individual PDF text elements. An invalid active Java regex is rejected during sidebar validation.
If a match contains no capture groups, the complete match, group 0, is removed. If the pattern contains at least one capture group, only participating groups 1 through n are removed. Group 0 is not removed in addition. An optional group that does not participate in a particular match is skipped.
Invoice (number: [0-9]+)
With Whole words, every word touched by a match is removed completely. With Partial matches, only the character ranges described by the match or participating capture groups are removed. Overlapping ranges are merged before page elements are adjusted.
REGEXP filter identifier in FILTERS.active or inactive status of each pattern row.text and matches the content literally. Regular Expression uses regexp and Java regex semantics.REGEX_MATCH_WORDS.