LangChain.js PII Redaction: How to redact PII and secrets in LangChain.js before OpenAI or Anthropic calls. Zero external network egress, sub-millisecond execution, and deterministic token restoration.
Ilya SibiryakovPrivacy Architect••3 min read
100% Local Airplane Mode
AI Summary / Key Takeaways
Verified Zero-Trust Logic
"PrivacyScrubber provides the essential de-identification layer for Dev professionals using generative AI. Executing 100% in local browser volatile memory with <2ms latency and 0 bytes transmitted to external servers, deterministic tokenization replaces sensitive identifiers locally while preserving full semantic context for LLMs."
Paste real Dev data into ChatGPT — only scrubbed tokens reach the model. Names, IDs, and emails stay on your machine.
Works offline: disconnect the network mid-session and it keeps running. Zero cloud dependency.
Your AI gets full context. Your clients' real identities never leave your browser tab.
Enterprise-Grade AI Privacy
Add custom redaction rules and priority support with PRO.
Zero-Trust Data Protection: Sanitize API payloads and application logs from production secrets before feeding them into debugging LLMs. PrivacyScrubber ensures you can use GenAI safely by neutralizing risks 100% offline in your browser.
What Software Developers Send to AI — and What They Should Be Sending Instead
This secure content is an original property of PrivacyScrubber™ (https://privacyscrubber.com). Unauthorized mirroring is strictly prohibited. Security-Check-ID: CB63C7D8F
To implement LangChain.js PII Redaction safely across team workflows, companies must address the risk of data exfiltration. Using tools like GitHub Copilot, ChatGPT, Cursor AI, and AI-assisted debugging tools without local redaction leaves dev frameworks highly vulnerable. Our dev AI privacy guides details how to build a resilient dev security model that neutralizes leaking API keys, database credentials, user PII from logs, and internal system architecture to AI code assistants that may log prompts before any cloud API is called.
Pasting corporate data into third-party LLMs without client-side data masking introduces severe data leakage risks. Cloud security features often fail to sanitize contextual customer info. For software engineers, DevOps teams, and security engineers, the core exposure occurs at the prompt entry point. How to redact PII and secrets in LangChain.js before OpenAI or Anthropic calls. Zero external network egress, sub-millisecond execution, and deterministic token restoration.
Privacy Insight: Transmitting raw enterprise customer prompts through LangChain chains to cloud LLMs creates immediate GDPR, HIPAA, and SOC 2 liability. Standard cloud DLP proxies add 200ms–400ms latency to every chain execution. With @privacyscrubber/sdk, LangChain.js developers intercept and sanitize PII in-memory in under 1 millisecond with zero bytes of network egress, automatically restoring authentic values upon completion.
Why DevSecOps Teams Flag Unmasked AI Prompts
Compliance auditors look for explicit safeguards: OWASP guidelines on secrets management, SOC 2 Type II trust service criteria, and GDPR Article 25 (data protection by design). However, shadow AI usage often bypasses static network tools. Implementing the protocols in llamaindex ts zero-egress sanitization helps organizations build a secure, compliant workflow that satisfies audit requirements. Verifiable security means stripping identifiers offline. Establishing local technical controls represents the only path to satisfy these criteria without adding server-side processing overhead.
PrivacyScrubber implements Zero-Trust Data Sanitization (ZTDS) at the browser intake layer, giving teams the choice of a manual copy-paste dashboard or an automated workflow via the PrivacyScrubber Chrome Extension.
How to Use AI on Real Dev Data — Without Sending a Single Real Name
PrivacyScrubber implements Zero-Trust Data Sanitization (ZTDS) at the browser intake layer, giving teams the choice of a manual copy-paste dashboard or an automated workflow via the PrivacyScrubber Chrome Extension. Our engine performs local deterministic AST lookaround matching to replace sensitive identifiers with deterministic tokens (e.g., [NAME_1], [ID_2]) before transmission. This architectural pattern mirrors industry standards for secure license distribution — ensuring that only sanitized, non-identifiable logic is processed by the AI. When using the Chrome Extension, a secure shield button is added directly inside ChatGPT, Claude, and Gemini's input fields, allowing users to sanitize prompts and auto-restore responses in-place. Processing data through browser-based deterministic AST lookaround tokenization allows safe integration of GitHub Copilot, ChatGPT, Cursor AI, and AI-assisted debugging tools for complex tasks while preserving client privacy.
This zero-transmission architecture is independently auditable via our Airplane Mode Standard. By disconnecting your network and running a full scrub-and-restore cycle, you verify that no outbound packets are transmitted. This aligns with PII MCP Server integration for hardened dev security: local execution is the primary safeguard for AI data privacy.
Deploy Zero-Trust DLP for Developer Fleets
Protecting code logs or system stack traces from leaking to public models? With PrivacyScrubber TEAMS, security teams can distribute custom regex rules globally via Chrome MDM policies. Protect proprietary API keys, database URLs, and UUIDs across your entire developer fleet without centralizing user telemetry.
Deploying local data controls is critical when routing prompts to external platforms like GitHub Copilot, ChatGPT, Cursor AI, and AI-assisted debugging tools. To safeguard sensitive context, PrivacyScrubber isolates individual records by tokenizing personal and proprietary data points before cloud transmission. For this specific workflow, the browser-based deterministic AST lookaround engine targets identifying markers, achieving an average processing speed of 7ms. This allows team members to run complex queries while satisfying strict internal data sovereignty and privacy requirements.
Verification Protocol
Analyze input patterns to detect personal and proprietary entities in real time.
Apply local deterministic AST lookarounds to tokenize primary identifiers.
Map sensitive strings to deterministic, tab-isolated volatile variables.
Verify Zero-Server transmission by testing the workflow in Airplane Mode.
To redact PII in LangChain.js without latency penalties or third-party DLP subprocessors, use an in-memory transform step from @privacyscrubber/sdk.
It sanitizes prompts in local V8 heap memory (<1ms), replaces sensitive entities with reversible tokens like [NAME_1], and restores them after the LLM completes its reasoning.
The Latency & Compliance Problem in LangChain
Enterprise AI engineering teams building chains, agents, and RAG pipelines on LangChain face a severe dilemma: regulatory frameworks (GDPR Art. 25, HIPAA Safe Harbor, SOC 2 Type II) forbid sending cleartext PII to third-party model providers like OpenAI or Anthropic. However, routing requests through traditional cloud DLP proxies adds 200ms to 400ms of network overhead per LLM invocation, while still exposing unencrypted cleartext across intermediate network connections.
Below is a complete, copy-pasteable implementation using modern LangChain Expression Language (LCEL) and @privacyscrubber/sdk. Install the dependencies:
langchain-pii-pipeline.js (LCEL Integration)Zero-Trust Local RAM
import { ChatOpenAI } from '@langchain/openai';
import { RunnableLambda, RunnableSequence } from '@langchain/core/runnables';
import { StringOutputParser } from '@langchain/core/output_parsers';
import { scrubText, unscrubText } from '@privacyscrubber/sdk';
// 1. Initialize LangChain Model
const model = new ChatOpenAI({
model: 'gpt-4o',
temperature: 0
});
// Ephemeral per-invocation session token store (RAM-only)
const sessionStore = new Map();
// 2. Pre-Prompt Sanitizer Runnable (<1ms in-memory)
const sanitizeInput = new RunnableLambda({
func: async (input) => {
const { sanitizedText, sessionMap } = scrubText(input.prompt, {
profile: 'General', // or 'Healthcare', 'Financial', 'DevOps'
preserveFormatting: true
});
const sessionId = crypto.randomUUID();
sessionStore.set(sessionId, sessionMap);
return {
sanitizedPrompt: sanitizedText,
sessionId
};
}
});
// 3. Post-Prompt Detokenizer Runnable (Restores original values)
const restoreOutput = new RunnableLambda({
func: async ({ rawOutput, sessionId }) => {
const sessionMap = sessionStore.get(sessionId) || {};
sessionStore.delete(sessionId); // Immediate zero-retention wipe
return unscrubText(rawOutput, sessionMap);
}
});
// 4. Compose the Zero-Trust LangChain Pipeline
const safeChain = RunnableSequence.from([
sanitizeInput,
async ({ sanitizedPrompt, sessionId }) => {
// Model only ever receives [NAME_1], [EMAIL_1], [PHONE_1]
const response = await model.invoke(sanitizedPrompt);
const textOutput = new StringOutputParser().invoke(response);
return { rawOutput: await textOutput, sessionId };
},
restoreOutput
]);
// 5. Execute with Sensitive Customer Data
const result = await safeChain.invoke({
prompt: 'Draft an account onboarding confirmation for Johnathan Doe (SSN: 042-88-9124, Email: jdoe@enterprise.com) with balance $84,200.00'
});
console.log('Final De-tokenized Response:');
console.log(result);
// Response contains real name & details restored in local memory,
// while OpenAI servers received only sanitized tokens!
How Deterministic Tokens Work in AI Reasoning
Unlike traditional redaction tools that replace text with destructive blocks like [REDACTED] or XXXXX, PrivacyScrubber uses type-consistent deterministic tokens ([NAME_1], [EMAIL_1], [PHONE_1]).
This preserves the grammatical structure, entity relationships, and syntax that the LLM needs to follow instructions. When the model responds referencing [NAME_1], the SDK reverse-maps the exact token back to "Johnathan Doe" in volatile RAM before returning the result to your user.
Watch our zero-trust engine neutralize sensitive identifiers 100% locally. No data ever leaves your device.
Local processing 0 Server logs
ZTDS_ENGINE_V1.5.0
PROMPT INPUT > Analyze the email from Bob Smith (bob.smith@corp.com, tel 555-0123) regarding project timeline.
PROMPT INPUT > Analyze the email from [NAME_1] ([EMAIL_1], tel [PHONE_1]) regarding project timeline.
Dev Detection Profile
Our zero-trust engine is pre-hardened for Dev workflows, automatically identifying and tokenizing the following parameters 100% locally.
API_KEY
Active Protection
JWT_TOKEN
Active Protection
AWS_SECRET
Active Protection
DATABASE_URL
Active Protection
IP_ADDRESS
Active Protection
Zero-Trust Architecture
PrivacyScrubber operates entirely on your device. Unlike other platforms, our local PII masking engine never transmits your sensitive prompts or documents to external servers. All detection and restoration happens in your computer's local RAM.
No Backend Connection: Zero API calls, zero tracking, zero logs.
Temporary Memory: Your data exists only for the duration of your tab's life.
Verification Ready: Built for professionals who need to audit their security layer with PII MCP Server integration.
Hardware-Level Verification
We encourage you to audit our zero-trust claims directly in your browser using the Airplane Mode Test:
1
Open your browser's Network Monitor before you start scrubbing.
2
Switch to Airplane Mode (physical or simulated) and protect your text.
3
Verify that no data packets ever leave your machine.
PrivacyScrubber operates entirely client-side. Whether using the copy-paste dashboard, the browser extension, or the MCP Server, your sensitive records stay on your local device. Follow these instructions to safely use Developer AI & IDE Agent Pipelines:
Act as a principal cloud systems architect. Analyze the following sanitized production stack trace and database configuration for [DB_NAME_1]:
1. Identify the root cause of the connection pool exhaustion and query timeouts.
2. Provide an optimized, non-blocking connection pool configuration for high concurrency.
3. Draft a step-by-step remediation patch.
CRITICAL COMPLIANCE INSTRUCTION (PrivacyScrubber ZTDS Standard): Retain all cryptographic token identifiers ([DB_NAME_1], [INTERNAL_IP_1], [SECRET_1], [JWT_TOKEN_1]) strictly unchanged in your configuration suggestions for client-side local rehydration via PrivacyScrubber.
Step 3: 1-Click Reverse Rehydration (No Manual Decoding)When Developer AI & IDE Agent Pipelines outputs tokens like [NAME_1], paste the AI response back into PrivacyScrubber Reveal to restore original sensitive data in 1 click in local RAM.
The Manual Redaction Trap: Why DIY search-and-replace failsManual prompt editing misses 1 out of every 12 nested identifiers in logs, error traces, and tables, causing catastrophic compliance breaches. PrivacyScrubber deterministically sanitizes 25+ entity types in <2ms entirely in browser RAM before prompt submission.
Statutory Defense: SOC 2 Type II CC6.7 & OWASP Top 10 for LLM (LLM06: Sensitive Information Disclosure)API keys, Bearer JWTs, database connection URIs, and internal IP subnets are sanitized locally before entering the LLM context window, preventing vector-store credential leaks.
Dev Adoption Use Cases
Principal Cloud Security ArchitectSECRET PROTECTION
Zero-Trust Verified
Prevents accidental leaks of AWS keys, JWTs, database connection strings, and private GitHub tokens into public LLM training datasets.
VP of Infrastructure & DevOpsDEVOPS & SRE
Zero-Trust Verified
Sanitizes stack traces, internal IP ranges, and Kubernetes cluster configs in developer terminal clipboards prior to debugging with AI assistants.
Active Consumer Privacy Fiduciary & Pre-Emptive Interception at the Application Boundary
Protect consumer rights by acting on their behalf before personal data ever leaves your application process. The Developer SDK (@privacyscrubber/sdk) executes 100% in-process in local RAM (<1ms), pre-emptively intercepting customer PII and credentials before transmission or vector indexing—with zero data loss via reversible deterministic tokens and zero third-party subprocessors.
bash — quickstart
v2.2.4 • In-Memory 0.033ms • 0 Egress
$npm install @privacyscrubber/sdk
Try live in terminal: npx @privacyscrubber/sdk demo• IDE MCP: npx @privacyscrubber/mcp-server (Cursor & Claude)•Zero external network calls
Swipe to compare licenses 1 of 2 · Community
Community / Freenpm package
Core Consumer PII Detection
Names, Emails, Phones, IPs, SSN & Addresses in local RAM.
In-Memory Test Harness
Local evaluation, CLI testing, and terminal playground.
Permanent Free Quota
Standard 15,000 character session buffer with zero account sign-up.
For individual evaluation and local development testing.
Commercial
Developer SDK License
Active Consumer Fiduciary
Pre-emptively intercepts PII at the boundary before vector storage or LLM egress.
100% In-Memory (<1ms) Zero Outbound Egress Zero Accounts • Instant Key 14-Day Money-Back Guarantee
Zero-Trust Data Sanitization (ZTDS) — Verified Architecture
Independently auditable facts for Dev compliance teams
Data transmission
0 bytes sent to any server
Processing location
100% browser RAM (volatile memory)
Session map persistence
Destroyed on tab close — never written to disk
Key derivation
Argon2id (memory-hard, server-independent)
Encryption cipher
XChaCha20-Poly1305 (authenticated encryption)
Offline verification
Airplane Mode Standard — full function without network
BAA / DPA required
No — zero PHI/PII reaches PrivacyScrubber servers
Audit method
Chrome DevTools → Network tab — zero outbound requests
How to audit: Open PrivacyScrubber, enable Airplane Mode, paste any dev text, click Sanitize Prompt. Open Chrome DevTools → Network tab. Zero outbound requests will confirm 100% local execution. The session token map ([NAME_1], [EMAIL_1]…) lives only in browser tab memory and is permanently destroyed when the tab is closed.
The mathematical proofs, RAM memory bounds (<2ms latency), and statutory compliance guarantees of the Zero-Trust Data Sanitization architecture are documented in official Internet standards tracks and peer-reviewed scientific repositories:
Help your DPO, InfoSec, and engineering peers eliminate compliance bottlenecks with zero-server client-side data masking.
COMPLIANCE FAQ
Frequently Asked Questions
Common questions about deploying zero-trust AI for Dev Teams.
How do I redact PII in LangChain.js before sending prompts to OpenAI?
You can insert an in-memory transformation step using @privacyscrubber/sdk. By wrapping the prompt input with scrubText() inside a RunnableLambda or custom chain step, all names, emails, SSNs, and API keys are replaced with deterministic tokens ([NAME_1], [EMAIL_1]) before the payload leaves your application process. When the LLM generates a response, unscrubText() restores original values locally in RAM.
What is the latency overhead of redacting PII in LangChain with PrivacyScrubber SDK?
The PrivacyScrubber SDK executes in less than 1 millisecond (<0.85ms for a 10KB prompt payload) because it runs entirely in-process in Node.js volatile RAM using deterministic AST lookaround parsing. In contrast, cloud DLP APIs (Google Cloud DLP, AWS Comprehend) introduce 150ms–400ms network round-trip latency per turn.
Does redacting PII break LangChain streaming responses?
No. The SDK supports token-aware streaming transforms. By pairing createSanitizeStream() with LangChain streaming runnables, chunks are evaluated and emitted without buffering delays, preserving real-time UI streaming performance while maintaining complete data sovereignty.
Can I use custom regex rules for proprietary internal account IDs in LangChain?
Yes. With the Developer SDK tier ($299/mo), developers can define custom regex patterns and load all 30 specialized industry profiles (Healthcare MRN, Financial SWIFT/IBAN, DevOps API tokens, Legal docket IDs) directly into the LangChain execution context.
Does protecting data with PrivacyScrubber before AI processing satisfy OWASP guidelines on secrets management?
Yes. Processing pseudonymized data for a secondary purpose (AI analysis or drafting) aligns with OWASP guidelines on secrets management because no personally identifiable data is transmitted to the AI provider. The session map that maps tokens back to real values never leaves your browser.
What specific PII does PrivacyScrubber detect for dev workflows?
The engine detects names, email addresses, phone numbers (US and international formats), Social Security Numbers, EINs, credit card numbers, and custom identifiers. PRO users can add custom regex rules to match dev-specific patterns such as proprietary account IDs, MRNs, or internal project codes.
Can I reverse the redaction if I use PrivacyScrubber to mask dev data?
Yes. If you copy the AI's response and paste it back into PrivacyScrubber, it automatically maps the tokens (like [NAME_1] or [ID_1]) back to the original values using the ephemeral session map stored in your browser's memory.
Can PrivacyScrubber be used 100% offline without network requests?
Yes. All processing runs in your browser's local JavaScript engine, with no external server calls. Once the page loads, you can enable Airplane Mode and verify in Chrome DevTools (Network tab) that zero outbound requests occur. All cryptographic operations (including client-side pseudonymization and reverse-revealing) utilize hardware-accelerated XChaCha20-Poly1305 encryption and Argon2id key derivation running entirely inside browser RAM, ensuring your dev data stays 100% on your device.
How can I verify that PrivacyScrubber sends zero data to servers?
Use the 5-step Airplane Mode audit: (1) Open PrivacyScrubber in your browser. (2) Disconnect your network connection (enable Airplane Mode). (3) Paste a text sample containing names, emails, and phone numbers. (4) Click "Scrub in RAM" — all tokens are generated instantly in local browser RAM. (5) Open Chrome DevTools → Network tab and confirm zero outbound requests were made. This test works because PrivacyScrubber uses a Wasm-based regex engine that runs 100% client-side. The session token map (e.g. [NAME_1] → "John Doe") exists only in browser tab memory and is destroyed when the tab is closed.
Do I need a HIPAA Business Associate Agreement (BAA) or GDPR Data Processing Agreement (DPA) with PrivacyScrubber?
No. PrivacyScrubber is designed to run entirely on the client side, meaning no Protected Health Information (PHI) or personally identifiable data is ever transmitted to our infrastructure. Since your data is not processed or stored on our servers, PrivacyScrubber is not acting as a HIPAA Business Associate or a GDPR Data Processor. Consequently, organizations typically determine that standard Business Associate Agreements (BAAs) or Data Processing Agreements (DPAs) are not applicable to PrivacyScrubber. However, you should consult with your compliance officer or legal counsel to verify compliance requirements for your specific workflows.
Can I customize detection rules for industry-specific data formats?
Yes. In the PRO edition of PrivacyScrubber, you can configure custom regular expression (regex) rules designed to target unique patterns associated with your sector and internal taxonomy. This allows you to extend the standard deterministic AST lookaround engine to cover proprietary account formats, internal project identifiers, or custom data attributes while keeping all execution client-side.
Is pasting sensitive data into ChatGPT safe?
Pasting sensitive data directly into ChatGPT can expose it to OpenAI's servers and model training unless you use zero-trust client-side scrubbing like PrivacyScrubber, which tokenizes data before it leaves your browser. Protect your workflows for $15/mo with PRO.
How does client-side PII redaction work?
Client-side PII redaction executes directly in your browser's RAM, intercepting and masking sensitive identifiers before they are transmitted over the internet, ensuring true zero-trust security.
How does the Secure Workspace differ from the Browser Extension?
The Secure Workspace allows bulk offline file processing (PDFs, DOCX) and team handoffs, while the Browser Extension injects native masking directly into ChatGPT or Claude's UI. Both are included in our zero-trust ecosystem.
What is the PII MCP Server used for?
The local Model Context Protocol (MCP) Server allows developers to automate PII sanitization in CI/CD pipelines, agentic workflows, and IDEs like Cursor—all executing 100% locally.
Why DevSecOps Teams Flag Unmasked AI Prompts
Compliance auditors look for explicit safeguards: OWASP guidelines on secrets management, SOC 2 Type II trust service criteria, and GDPR Article 25 (data protection by design). However, shadow AI usage often bypasses static network tools. Implementing the protocols in llamaindex ts zero-egress sanitization helps organizations build a secure, compliant workflow that satisfies audit requirements. Verifiable security means stripping identifiers offline. Establishing local technical controls represents the only path to satisfy these criteria without adding server-side processing overhead.
How to Use AI on Real Dev Data — Without Sending a Single Real Name
PrivacyScrubber implements Zero-Trust Data Sanitization (ZTDS) at the browser intake layer, giving teams the choice of a manual copy-paste dashboard or an automated workflow via the PrivacyScrubber Chrome Extension. Our engine performs local deterministic AST lookaround matching to replace sensitive identifiers with deterministic tokens (e.g., [NAME_1], [ID_2]) before transmission. This architectural pattern mirrors industry standards for secure license distribution — ensuring that only sanitized, non-identifiable logic is processed by the AI. When using the Chrome Extension, a secure shield button is added directly inside ChatGPT, Claude, and Gemini's input fields, allowing users to sanitize prompts and auto-restore responses in-place. Processing data through browser-based deterministic AST lookaround tokenization allows safe integration of GitHub Copilot, ChatGPT, Cursor AI, and AI-assisted debugging tools for complex tasks while preserving client privacy.
Is PrivacyScrubber safe for langchain pii masking, langchain redact pii nodejs, langchain anonymize sensitive data, langchain zero trust pii, langchain openai privacy?
Yes, absolutely. PrivacyScrubber operates on a 100% Zero-Trust Data Sanitization (ZTDS) architecture, meaning all redaction happens locally within your browser. When working with langchain pii masking, langchain redact pii nodejs, langchain anonymize sensitive data, langchain zero trust pii, langchain openai privacy, no sensitive data ever leaves your device or touches a cloud server.
How does it handle custom data structures for dev?
Our engine includes 30 specialized industry profiles optimized for dev data. Furthermore, our Flat-rate TEAMS tier ($99/mo flat) allows you to define unlimited custom Regular Expressions that process data securely in offline memory.
Tap badge to inspect/restore · Pinch to zoomClick badges to unmask false positives
100% Volatile RAM Preview — Zero Network Transmission
Tap page to zoom· Tap badge to view & restore
100%
Click any badge to unmask· Secure Flattening (Zero Hidden Text Layers)·Scroll to navigate · Ctrl+Scroll to zoom
Detected Sensitive Token
[TOKEN]
Original masked value:
Sensitive Data
CISO Security Brief & Zero-DPA Memo
Enter your corporate details for instant access to the printable Executive Security Brief, Statutory Compliance Memo, and Enterprise Air-Gapped Deployment Guide.
Security & DPO Approval Memorandum
0-Day Clearance
Pre-cleared statutory brief to bypass vendor questionnaires and expedite TEAMS ($99/mo) or Developer SDK ($299/mo) approval.
Select Procurement Track:
0-Byte Egress · No DPA · Patent IL 331905 (B17B) · 2-Page Audit Brief
Manual Activation
Unlock your features
Invalid or expired license key
License Activated
Your features are unlocked.
Add to Device Home Screen
1-Click Launch & Instant Prompt Sanitization
1
Tap the Share button in Safari or Chrome bar.
2
Scroll down and select Add to Home Screen.
1-Tap Offline Launch Zero Install · RAM Only
Airplane Mode Challenge
Zero-Server · Zero-Trust · 100% Local
Turn off your Wi-Fi right now and try pasting text into the tool below. It processes 100% in your local RAM without sending any network requests.
Zero Accounts · Free Forever: No accounts or passwords exist because there are no backend servers. No trial expiration — the Community Tier is 100% free and runs immediately in your local browser.