At a high-level, the workflow schedule was this,
- Client solicitors sent an email to a specific Outlook address requesting a search of the client’s details in the database. For the search, we needed the first and last name of the client, together with their date of birth and current address
- A search could only proceed if there were corroborating documents attached to the email to “prove” who the client was.
- The attachments would typically be PDFs containing images of passports, driving licences, and other official documents. In addition, we could also receive MS Word documents or PDFs of bank statements and utility bills
- At a specified time each day, the IDP processed these emails, extracted and classified the attachments and extracted any relevant Personally Identifiable Information (PII), which was sent to a claims database for searching.
- If claim details were found in the database, a password-protected PDF containing those details was created and sent back to the solicitor, along with a separate email containing the password to unlock it. If no results were found, a simple email was sent back indicating this.
Apart from a Human-In-The-Loop (HIL) process that allowed users to perform a final cross-check as to whether the PII extracted from different documents matched, the process was fully automated.
Incidentally, this new process proved approximately 90% more efficient than the manual system it replaced.
In this article, I’m going to show you the steps I took to implement this system on the AWS cloud platform, albeit a slightly simplified version of it.
Note that apart from being a user of Amazon Web Services (AWS), I have no affiliation or association with the company.
In the real system I developed, our source of emails came from a corporate Outlook email account. I don’t have access to one of those now, so instead, I’ll use my personal Gmail as our email server and assume each email contains only one attachment. I’m assuming that any attachment will always be an image file (JPG or PNG) of a passport, driving licence, or bank statement. Also, the final step in our process will be a simple result file on S3 that indicates the classification and PII extraction values for the attachment.
Due to the length of the code snippets etc., I’ll include a link at the end of the article to a GitHub repo, where you can find all the Lambda code, the Step function code, IAM permissions, and everything else required for this project.
As this is mainly about the AWS side of things, I’m going to assume that everything is set up in your email client to enable an AWS Lambda process to read your emails.
For Gmail, briefly, this process would include:-
- creating or selecting a Google Cloud project
- enabling the Gmail API
- enabling OAuth and completing a one-time Google consent flow.
After that, you would store the returned credentials/refresh tokens securely in something like AWS Secrets Manager.
The goal is to extract the following pieces of PII from each email attachment.
Document Type, First Name, Last Name, DoB and Address
Our test Gmail attachment images
I asked Codex to create 3 obviously fake images of a passport, a driver’s licence, and a bank statement. Although fake, they contain the same information as the real thing. Here are those images.
I emailed these images to my Gmail account and labelled the emails “DocumentProcessing”.
We now have everything in place to start creating our IDP.
Here are the AWS services we’ll use to implement our IDP.
- IAM for permissions
- EventBridge for scheduling
- Step function for orchestration
- Lambda for compute
- S3 for storage
- Secrets Manager for storing OAuth and other secret information
- Textract for OCR
- Bedrock for AI inference
1/ Storing our OAUTH credentials securely
The first thing we need to do is store our Gmail OAuth credentials and refresh token in AWS Secrets Manager. For me, my email setup returned a JSON document that looked like this.
{
"installed": {
"client_id": "my_client_id.apps.googleusercontent.com",
"client_secret": "MY_CLIENT_SECRET",
"project_id": "my_project_id",
"auth_uri": "https://accounts.google.com/o/oauth2/auth",
"token_uri": "https://oauth2.googleapis.com/token"
}
}
Open the AWS management console and go to the Secret Manager screen. Click the “Store a new secret” button and choose “Other type of secret” as the secret type. Under the key/values pairs section, select the Plaintext TAB and enter your JSON from above. So your screen should look like this,
Click Next, then enter a suitable name and description for your new secret. Click Next, then Next again, before clicking the Store button.
2/ Set up our S3 buckets
We’ll have one general-purpose bucket and 5 folders that will hold our data.
- raw: this is where the raw email attachments will be downloaded to
- textract: this will hold the text returned by Textract
- results: this will hold the results of the document classification and PII information
- audit: logging information
- quarantine: any unclassified images/documents will be placed here
3/ Creating our Lambdas
We’ll need three.
i) ingest_email_image retrieves eligible email messages and writes each image to our raw folder on Amazon S3.
ii) extract_text sends the S3 images to the Amazon Textract service and stores the extracted text.
iii) classify_and_extract_pii sends that extracted text to Amazon Bedrock Claude Sonnet 4.6, which attempts to classify the document type and extract any relevant PII and writes the result.
Here’s a simplified diagram of the process we’ll build.
EventBridge Scheduler
|
v
Standard Step Functions workflow
|
+-- Ingest Gmail image attachments -> S3 raw/
|
+-- Map: one email/image at a time
|
+-- Textract Lambda -> S3 textract/
|
+-- Bedrock Lambda -> S3 results/
|
+-- Record a success or failure result
Lambda 1 – ingest_email_image
The first Lambda polls our configured Gmail label using read-only OAuth credentials from AWS Secrets Manager. It validates that each email contains exactly one JPEG or PNG attachment, saves valid images under the S3 raw/ folder, and returns outputs for the next processing stage; invalid or duplicate attachments are reported separately.
After this Lambda runs, you should see the three image files saved under the raw folder of your S3 bucket.
Also note that the Lambda uses the following environment variables.
Variable Value
-------- ------
DOCUMENT_BUCKET YOUR_OUPUT_BUCKET_NAME
GMAIL_LABEL_NAME YOUR_GMAIl_LABEL_NAME
GMAIL_SECRET_ARN YOUR_SECRETS_MANAGER_ARN
LOG_LEVEL INFO
MAX_IMAGE_BYTES 10000000
MAX_MESSAGES_PER_RUN 25
Lambda 2 – extract_text
The second Lambda takes the images stored under raw/ and runs Amazon Textract’s synchronous document-text detection on them. It saves the extracted text, individual lines, confidence scores, and source metadata as JSON under the S3 textract/ folder,
Here is a prettified, shortened version of the typical output when I ran this against the bank statement image.
{
"schema_version":1,
"processed_at":"2026-06-25T11:04:19.850574+00:00",
"source":{
"bucket":"MY_BUCKET",
"image_key":"raw/bank.png",
"gmail_message_id":"19efe3cd80025b99",
"gmail_attachment_id":"ANGjdJ-NRkTD-45NFgNRZfHyWSKezw5MI03xN7bvTs4PZKuDjbMCZPk--bnSp4lMVldYhISyd64jH9JgN5cJXTm2_KVDFDsrIFDWJAWs7K6CtMJvsb9OYNX7GZSnwptK_U6M3npTu_02mNbOY_Q-PTKt58rcYbIOA7aDbOxlrAqFh3NmoT69FV92oMjy6NqaT1uSXXZY40h62Ofp-JT4s2JGLKCymhi_o49eAPvBj8azguB4WyoowqWiQEh_cL5IdI7uJC7GEIu9LoUZYuxr7AHZBwACeasd-sNbKLYUb7FwDhPpla_iGD9aDyncH8Q",
"mime_type":"image/png"
},
"textract":{
"model_version":"1.0",
"document_metadata":{
"Pages":1
},
"lines":[
{
"text":"UK Bank Statement",
"confidence":99.98
},
{
"text":"bank",
"confidence":99.99
},
{
"text":"Unastign Uiver Pol, 01426",
"confidence":93.85
},
...
...
...
{
"text":"Terms UK Banks statement fan Their Band condition, dlesstery. AT stalewaly for diescavey fevien 720, 00553 Dear",
"confidence":93.75
},
{
"text":"Contrined Parih Bank ellamty of allimd elines and ptke and corv. llartes hier. uniedarfc lard sand colunrty mp conmontivre",
"confidence":79.02
},
{
"text":"tharter m allar mpercian the Cant'g paymer Is amowe Sane UK Banklila the Serd Promonry à a corditions of larril and",
"confidence":78.91
},
{
"text":"Statemant of Teard of thre D2. Carpltionrnes Other whitan and Contiiomed Earm Conditions of Terms witte pittier tied Oswllas",
"confidence":89.62
},
{
"text":"Terms and Condition, lod Cporl and Balance.",
"confidence":82.19
}
],
"text":"UK Bank Statement\nbank\nUnastign Uiver Pol, 01426\nUttesline 952016\nVobue. 4346\nContasc! 84 6207\nContasc! 16777982\nMame Holder\nAccount\nJohn Smith\nJohn Smith\n11 The High Street\nNewTown\nNT1 3WE\nAccount NU83632\nNome\nStatement Period\nSort Code\nValid\n01/05/2026 to 01/06/2026\nStatement Period\nOpening Balance\n01/05/2026 to 01/06/2026\nClosing Balance\nTransactions\nto\nDate\nDescription\nAmount\nBalance\n07906TN\nDirect Debits\n£72,000\n£59,000\n080161B\nCard Payments\n£73,000\n£95,000\nCard Payndit\n090061N\nSalary Credit\n£95,000\n£56,000\n064061B\nSalary Credit\n£23,000\n£64,000\n095061D\nATM Withdrawals\n£60,000\n£60,000\n075061N\nATM With drawaly\n£60,000\n£64,000\n374091D\nATM Vithdrawaly\n£60,000\n£63,000\n340061D\nATM Withdrawals\n£60,000\n£73,000\n153381D\nATM Withdrawaly\n£60,000\n£13,000\n150491D\nATM Withdrawaly\n£10,000\n£13,000\n169010P\nATM Withdrawals\n£11,000\n£19,000\n188191D\nATM Withdrawals\n£10,000\n£95,000\n11111TD\nSalary Credit\n£60,000\n£55,000\n11211TD\nATM Withdrawals\n£60,000\n£95,000\nPage Number\nTerms UK Banks statement fan Their Band condition, dlesstery. AT stalewaly for diescavey fevien 720, 00553 Dear\nContrined Parih Bank ellamty of allimd elines and ptke and corv. llartes hier. uniedarfc lard sand colunrty mp conmontivre\ntharter m allar mpercian the Cant'g paymer Is amowe Sane UK Banklila the Serd Promonry à a corditions of larril and\nStatemant of Teard of thre D2. Carpltionrnes Other whitan and Contiiomed Earm Conditions of Terms witte pittier tied Oswllas\nTerms and Condition, lod Cporl and Balance."
}
}
Lambda 3 – classify_and_extract_pii
The final Lambda takes each block of JSON text that Textract scraped from the input images and uses AWS Bedrock to classify the document type and extract any PII it can glean. The key to success in this step is the model used and the prompt passed to it. For my work project, I used Claude Sonnet 4.6 and a very large prompt because it had to handle more complex inputs. But for this more straightforward example, while we’ll still use Sonnet 4.6, we can get away with a much simpler prompt.
First, we need to note the model ARN associated with Sonnet 4.6 in the region you’re working in. You do this from the Bedrock console, so select that, then click on the Inference Profiles link on the left-hand menu bar. Search for Sonnet 4.6 and note the appropriate ARN. As this is an inference profile, you will also need to take note of the regions that AWS can route inference calls through, as these will need to be added to your IAM permissions together with the inference profile permission.
Here are the final outputs I received after processing the bank, passport and driver’s licence attachments.
Bank attachment
{
"schema_version":1,
"processed_at":"2026-06-25T11:04:23.221887+00:00",
"source":{
"bucket":"MY_BUCKET",
"textract_key":"textract/bank.png.json"
},
"model":{
"model_id":"us.anthropic.claude-sonnet-4-6"
},
"classification":"bank_statement",
"pii":{
"first_name":"John",
"last_name":"Smith",
"date_of_birth":null,
"address":"11 The High Street, NewTown, NT1 3WE"
}
}
Passport attachment
{
"schema_version":1,
"processed_at":"2026-06-25T11:04:21.829715+00:00",
"source":{
"bucket":"MY_BUCKET",
"textract_key":"textract/passport.jpg.json"
},
"model":{
"model_id":"us.anthropic.claude-sonnet-4-6"
},
"classification":"passport",
"pii":{
"first_name":"AVERY",
"last_name":"EXAMPLE",
"date_of_birth":"1991-08-30",
"address":"14 Example Lane, Testford, TE1 2ST"
}
}
Driver’s licence attachment
{
"schema_version":1,
"processed_at":"2026-06-25T11:04:22.507427+00:00",
"source":{
"bucket":"MY_BUCKET",
"textract_key":"textract/driving.jpg.json"
},
"model":{
"model_id":"us.anthropic.claude-sonnet-4-6"
},
"classification":"driving_licence",
"pii":{
"first_name":"AVERY",
"last_name":"EXAMPLE",
"date_of_birth":"1991-08-30",
"address":"14 Example Lane, Testford, TE1 2ST"
}
}
This Lambda uses the following environment variables.
Variable Value
-------- ------
BEDROCK_MODEL_ID us.anthropic.claude-sonnet-4-6
DOCUMENT_BUCKET MY_BUCKET
LOG_LEVEL INFO
MAX_INPUT_CHARS 30000
MAX_JOBS_PER_RUN 25
The Step Function
Although for this simple example we could probably just daisy-chain the three Lambdas to run one after another in code, the best practice is to use AWS’s orchestration tool, called Step. Step functions have the added advantage of being retryable after failures and elegantly handling errors and timeouts.
AWS uses a language called the Amazon State Language (ASL) to define their Step Functions. The ASL code is in the repo. Within Step, you can render a process into a useful flow diagram, which, in our case, looks like this.
It might look a bit complicated, but a lot of that is the typical scaffolding that surrounds a process like this such as error handling. It’s basically just orchestrating the three Lambdas by calling them one after the other.
The EventBridge scheduler
To tie everything together, I needed a way to kick off the Step function on a daily basis. An easy way to do that on AWS is via a service called EventBridge. EventBridge is capable of many things, but one of its most useful functions is to set up a cron based timing event that can call other AWS systems at a specified date and time. On my work project, I set a scheduled task to run at 11:55 PM each night to catch all emails in the inbox for that day, so I’ll repeat that here.
Click on the EventBridge service from the main console, then click the Schedule menu item on the left-hand bar. On the screen display, click the “Create schedule” button. You should see a screen like this
Type in a name and description for your schedule. Leave the “Schedule group” field as the default, then choose the “Recurring schedule” option. Select your required Timezone and opt for the cron-based schedule. From here, fill out the fields as you would for a cron job.
For our example of Monday through Friday at 11:55 PM, this would be
Next, set a flexible time window if you want, along with optional start/end dates. Click the Next button and select “AWS Step Functions” as your target. Select the Step function name you want to run from the drop-down list and enter any inputs.
Click Next. On this screen, enter NONE for the action to take after the event has finished. You can also enter a dead letter queue (DLQ), encryption details, and IAM permissions.
Click Next once again to go to the Review screen, and if you’re happy with all the details, click the “Create schedule” button.
Your schedule will now be “live”, and your Step function will execute every weekday at 11:55 PM.
What’s needed to turn this into a production system?
So, what I’ve shown you in this article is the bare bones of an automated IDP system. Looking back at the real-life IDP system I implemented, here are the changes you would need to properly productionise this.
- Handle more types of attachment documents, e.g., MS Word, plain text, PDF, images, and handwritten material.
- Handle multiple attachments per email.
- Code a front-end HIL so that users can cross-check the returned PII from different documents to ensure they match.
- Have the HIL front end automatically trigger the next automated stage, e.g. a separate Step function that implements …
- Sending the PII to the back-end database for search.
- Automating the production of PDFs containing database search results.
- Automatically password-protecting the PDFs containing the database search results.
-
Automating the sending of the PDFs and the separate password file back to the requester.
-
Ensuring everythingis logged
Note that there is a cost associated with building and developing the system I’ve described above so if you do follow along and create real resources on AWS, please be aware of this and delete anything you no longer require to avoid surprise charges.
Summary
This article walks through the process of building a simplified Intelligent Document Processing system on AWS. The system starts with emails with a specific label in Gmail, extracts image attachments to Amazon S3, uses Amazon Textract to read the text, and then sends that text to Amazon Bedrock with Claude Sonnet to classify the document and extract key PII such as name, date of birth, and address. This work is carried out by AWS Lambda functions.
The Lambda calls are orchestrated with a Step Function, scheduled by EventBridge, and secured with IAM permissions and Secrets Manager. The result is a practical AWS pipeline that demonstrates how email-based document intake, classification and data extraction can be automated while remaining production-aware in design.
We covered a lot of ground and probably over-engineered the solution based on the “simplified” input images we were trying to process. For our test cases, we could probably have bypassed the AWS Textract step altogether and just had Bedrock classify and extract the information directly from the images.
I needed the Textract step in my work project because the input I was dealing with was more complex, including PDFs and images of hand-written text. I think it was worth keeping in that step to show that you have options if your inputs are less than straightforward.
Finally, I mentioned ed some steps you would need to take to turn the system I described into a proper, production-ready process.
For all the code and ancillary files, check out my GitHub repo at the link below. Note that I have redacted ARNs, bucket names, AWS account numbers, and anything else that may pose a privacy or security risk.
PS I’m in the market for contract work just now. If you or someone you know is looking for an experienced data engineer, either remote or Edinburgh, UK-based, with skills in AWS, AI, Python, SQL, PySpark, DuckDB, etc., let me know. You can find me on LinkedIn.