Patterns & Sets in PII Discovery and Data Synthesis
IRI Workbench provides two complementary resources for sensitive data discovery and data synthesis: Regular Expression (RegEx) patterns and set files. Patterns identify PII and other sensitive values by structure. Set files support dictionary-based matching and provide realistic values for pseudonymization and test data synthesis.
Success in data governance and application development often depends on two related workflows. First, you identify sensitive information. Then, you manage that information responsibly. Without the right tools and techniques, finding PII and transforming it into usable, safe data can become labor-intensive and manual.
The IRI Workbench graphical user interface (GUI) supports sensitive data discovery, value replacement (pseudonymization), and test data synthesis. It provides these capabilities for IRI Data Protector Suite tools such as FieldShield, DarkShield, and RowGen. IRI Workbench also provides two professionally vetted, out-of-the-box resources for these workflows: Patterns and Set Files.
RegEx patterns and Set File values help identify sensitive data. Set files also support data replacement and value selection. Together, these resources help streamline data discovery, masking, and DevOps processes.
Data Discovery via Patterns
To identify sensitive data, IRI Workbench includes an extensive out-of-the-box library of Regular Expression (RegEx) patterns. These patterns can scan data sources for PII and other structured information, including government identifiers, license plates, financial identifiers, and timestamps.
The pattern list shown in the tables below, but not the actual expressions themselves, identifies the RegEx expressions that IRI provides out of the box for FieldShield, DarkShield, and CellShield EE users during data discovery. Users can apply these patterns to column names as location matchers or to actual values in columns or files as data matchers.
Data Discovery via Set Files
Patterns are not the only way to recognize sensitive information. During data classification searches through relational databases or unstructured documents, the Set File Matcher can identify values by performing a dictionary lookup against a supplied text file containing strings.
DarkShield, for example, can scan content for either:
- Literal matches — exact matches to values in a predefined set file.
- Fuzzy matches — close matches to values in that set file.
Patterns and set files provide complementary approaches to discovery: RegEx patterns identify data based on defined patterns, while set files identify data by matching against predefined values.
Pseudonymization via Set Files
Once you identify sensitive data, you need a way to replace it while preserving the utility of your datasets.
IRI set files can supply replacement values for this purpose. You can use libraries containing names, demographic information, regional identifiers, and other real-world values to pseudonymize production information with realistic alternatives.
For example, you can replace a production name with another plausible name drawn from an appropriate set file rather than with an arbitrary string. This helps preserve the utility of the transformed data while replacing the original sensitive value.
Test Data Synthesis via Set Files
IRI RowGen also uses set files as repositories of real-world values for test data synthesis. Values can be randomly selected from those files to generate semantically realistic data such as:
- Names
- Addresses
- Religions
- Diagnoses
- Departments
RowGen can write the generated values into target database columns, flat files, and reports. To maintain referential integrity across related database tables, RowGen can automatically point parent- and child-table generation rules to the same set file, ensuring that selected values remain consistent across the database.
Set files can therefore be used across several workflows, including sensitive data discovery, pseudonymization, and test data synthesis. Whether the objective is to search for PII, pseudonymize production data for privacy-law compliance, or create entirely new data for application development, set files provide reusable source values that help keep the resulting data realistic and functional.
Provided Patterns in IRI Workbench
IRI Workbench includes an extensive library of out-of-the-box RegEx patterns for use in data discovery. Each copy of IRI Workbench includes these patterns, and users can modify them or introduce new ones. Users can also enhance pattern matching with computation verification through data class validators.
The complete library spans country-specific and global formats. Because the supplied catalog is extensive, the examples below illustrate the types of patterns included rather than reproducing every individual expression or example value.
Africa
The provided African resources include, for example, a Nigerian license plate pattern.

Americas
The Americas library covers a broad range of personal, government, financial, health, and location identifiers.

Asia
The pattern library includes resources for many Asian and Middle Eastern formats.

Europe
IRI Workbench also provides an extensive range of European patterns.



Oceania
Australian patterns include examples for mobile and telephone numbers, PINs, postal codes, state names, tax values, passports, bank accounts, driver’s licenses, Medicare and health insurance identifiers, national telephone numbers, Tax File Numbers, Business Numbers, and Company Numbers.

Global Patterns
The global pattern library covers values that are not limited to one country. The catalog includes multiple patterns for common credit-card formats, including American Express, China UnionPay, Diners Club, Discover, JCB, Mastercard, and Visa. Additional global patterns include YYYYMMDD date formats, passwords of 8–20 characters, international phone formats, international securities identifying numbers, and additional payment-card formats.

Provided Data Sets (Set Files)
The diagrams below show the set files IRI ships with each deliverable. A primary category consists of people’s names, with values drawn from researched lists of female and male first names by country, along with last (family) names. These set files can be used for PII discovery, pseudonymization, and test data synthesis.
Country-Based Name Sets
IRI also provides the same types of name set files, including female and male first names and last names, for countries throughout Asia and Europe, as shown below.

US Set Files
The root US folder contains larger first- and last-name collections, along with additional set files for healthcare, locations, names, and other data used in discovery, masking, and test data projects.

International and Miscellaneous Set Files
IRI Workbench also includes an international miscellaneous folder containing set files that can be used in other projects, including PII discovery for GDPR compliance.

These examples demonstrate that set files are not limited to names. They can provide reusable collections of values for multiple types of discovery, masking, and data-synthesis jobs.
Custom Patterns and Set Files
The RegEx patterns provided in IRI Workbench can be modified, and users can introduce new patterns as needed. Pattern matching can also be enhanced through data class validators.
Just as you can create new patterns, you can also create and provide your own set files for use in these jobs. See this article for an example of creating set files in IRI Workbench.
Summary
IRI Workbench provides an extensive library of RegEx patterns and set files for PII discovery, pseudonymization, and test data synthesis. RegEx patterns support sensitive data discovery by matching defined data formats, while set files can identify values through dictionary matching and provide realistic replacement or generated values for masking and test data workflows.
These resources support IRI technologies including FieldShield, DarkShield, CellShield EE, and RowGen. IRI Workbench includes country-specific and global patterns, name and other reusable set files, and the ability to modify existing patterns or create custom patterns and set files for specific requirements.
Using IRI’s provided RegEx patterns and set files can reduce the time required to research and assemble these resources. They can also be used without exposing data or systems to cloud or AI services that may not provide the security or reliability required.
Frequently Asked Questions
What is the difference between a pattern and a set file in IRI Workbench?
A pattern uses a Regular Expression to recognize data according to its structure or format. A set file contains known values that can be compared against data through dictionary matching and can also provide replacement or generated values for pseudonymization and test-data synthesis.
Can DarkShield use set files for both exact and approximate matching?
Yes. The Set File Matcher can identify sensitive information through a dictionary lookup, and DarkShield can scan for either literal (exact) or fuzzy (close) matches against a predefined list.
How does RowGen use set files for test data?
RowGen can randomly select real-world values stored in set files and write them into target database columns, flat files, and reports. These values can represent items such as names, addresses, religions, diagnoses, or departments.
How does RowGen maintain referential integrity when using set files?
For related parent and child tables, RowGen can automatically point the generation rules to the same set file so that selected values remain consistent across the database.
Can users create their own patterns and set files?
Yes. IRI Workbench allows supplied patterns to be modified and new ones to be introduced. Users can also create and provide their own set files for applicable jobs.










