Virtusa logo

GEN AI Data Architect

Virtusa

RemoteFull timeMid levelPosted today
Apply with JobAssist

About the role

Job Requirements Role summary

Create high-quality, customer-specific synthetic data and own RAG / knowledge pipelines so each deployment of

CCAI, voice, and chat can be configured, grounded, demonstrated, and validated without using real customer PII.

You design generation and ingestion pipelines and load data into the correct GCP and AWS services.

What success looks like

Each customer engagement has a documented synthetic dataset covering the channels in scope

  • Each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable,

and evaluated.

Data and retrieval quality are good enough for configuration, evaluation, and stakeholder demos, and

safe enough for isolation and compliance expectations.

Generation and indexing are parameterized and repeatable, not a one-off manual copy-paste per

customer.

Key responsibilities

Analyze each customer’s domain: intents, entities, knowledge topics, document types, languages, tone,

and edge cases.

Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.

Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structured

entity values.

Schedule and document index refresh processes when customer knowledge changes.

Use appropriate techniques while

documenting parameters and limitations.

Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and source

corpora.

Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized per

customer.

Partner with the Conversational Platform Specialist so loaded data and indexes actually drive the

deployed experience.

Partner with DevOps so pipeline jobs, stores, and secrets are automated and isolated per customer.

Required Qualifications 4+ years in data engineering, conversation design operations, applied NLP data work, or knowledge-

pipeline engineering.

Working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-

grounded knowledge data.

Strong judgment on synthetic-data quality, retrieval quality, and privacy safety.

Preferred Qualifications LLM-assisted synthetic data generation in a production or implementation setting.

Familiarity with BigQuery, S3, and document stores used as knowledge sources.

Multilingual data generation or evaluation experience.

Work Experience 7-10Years

Millions of jobs, with real people getting hired every day

20,000+
New jobs added daily
7,000,000+
Verified job listings
500,000+
Tailored applications submitted
FAQ

Questions, answered

Click "Apply with JobAssist" – we tailor your resume and application to this role and submit it for your approval.

Yes. This role at Virtusa was screened before publishing – we confirmed the employer before listing it.

The employer didn't disclose a salary range for this listing. JobAssist shows pay whenever it's available.

This position can be done from anywhere, with no in-office requirement.

Yes – every application is tailored from your profile and this job's requirements, and you can review and edit before it's sent.