Role: Prompt Engineer
Location: 210 Hudson Street, Jersey City, NJ, 07311
Duration :- Long Term Project
(candidates must be willing to work onsite 4 days a week at the client’s NJ office)There may also be a requirement for an in-person (F2F)
LLM Interaction Design / Prompt Optimization / GenAI Application Quality
Level
Specialist Individual Contributor
Target / alternate titles
Prompt Engineer; LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist
Core keywords
prompt engineering, system prompt, few-shot, RAG, prompt evaluation, prompt injection, jailbreak, conversational AI, LangChain, Semantic Kernel, LlamaIndex, AWS Bedrock, Azure OpenAI, Copilot Studio, Power Platform
Role purpose
Design, test, govern, and continuously improve prompts, system instructions, conversation flows, and interaction patterns for AIRP LLM applications and related citizen-development experiences. The role ensures model outputs are accurate, grounded, safe, consistent, cost-aware, and aligned with business and compliance expectations.
Client-specific emphasis
· Prompt work must support enterprise business use cases, not generic chatbot experimentation.
· Reusable prompt patterns should be suitable for AIRP and, where applicable, Copilot Studio / Power Platform citizen-development scenarios.
· Candidates must understand prompt security, sensitive data handling, citations/grounding, and structured evaluation.
Primary ownership
· Prompt patterns, system instructions, response templates, and conversation policies for AIRP LLM use cases.
· Prompt testing, versioning, evaluation, and quality-improvement workflows.
· Reusable prompt libraries and guardrail patterns for business teams and responsible citizen development where applicable.
Key responsibilities
· Design prompts for chatbots, copilots, RAG systems, document analysis, summarization, workflow agents, knowledge assistants, and decision-support experiences.
· Develop system prompts, few-shot examples, tool-use instructions, response formats, escalation logic, citation behavior, and conversation policies.
· Optimize prompts for KYC support, credit underwriting support, governance tracking, pitch book generation, Banker 360, Customer 360, deal library intelligence, financial crime quality, and sanctions screening use cases.
· Build reusable prompt libraries and templates aligned to enterprise standards, business domains, AIRP patterns, and citizen-development guardrails.
· Evaluate prompt performance using metrics such as task success, groundedness, hallucination rate, completeness, safety, user satisfaction, latency, and token cost.
· Partner with engineers to implement prompt versioning, testing, deployment, and monitoring in production systems and CI/CD workflows.
· Support RAG quality by assessing retrieval context, chunking quality, source citation behavior, response synthesis, and missing-context behavior.
· Conduct adversarial testing for prompt injection, jailbreaks, instruction conflicts, sensitive-data leakage, unsafe outputs, and unauthorized tool use.
Must-have candidate profile
· Strong understanding of LLM behavior, prompt design, tokenization, context windows, RAG, embeddings, and model limitations.
· Hands-on experience with OpenAI APIs, Azure OpenAI, AWS Bedrock, Anthropic, LangChain, LlamaIndex, Semantic Kernel, Copilot Studio, or similar platforms.
· Ability to debug LLM outputs using structured testing, error analysis, and iterative refinement.
· Strong writing, analytical, communication, and stakeholder-management skills.
· Understanding of prompt-security risks including prompt injection, jailbreaks, data leakage, hallucination, and instruction conflicts.
· Ability to create repeatable prompt templates and evaluation evidence suitable for enterprise governance.
Preferred experience
· Background in NLP, conversational AI, UX writing, technical writing, product design, knowledge management, business analysis, or financial-services operations.
· Experience in financial services, legal, compliance, risk, operations, customer support, banker productivity, or enterprise knowledge domains.
· Familiarity with Microsoft Copilot Studio, Power Platform, prompt registries, A/B testing, human review workflows, and evaluation tooling.