Search

Sr AI Engineer - OCR

PublishedPublished: 6/14/2022
Technology

Job Description

We are looking for 5+ years of relevant experience and: Strong OCR experience across LLMs, cloud tools, and traditional tools such as Tesseract Experience with handwriting, poor-quality faxes, image artifacts, and database validation Understanding of current tool/data limitations and what is required to improve quality Strong Python and SQL; familiarity with GCP Experience developing, deploying, and maintaining production applications Strong business framing, stakeholder communication, and technical judgment Demonstrated end-to-end ownership of ambiguous projects, with limited supervision An outcome-oriented mindset focused on solving the business problem, not just executing a task Successful candidates should be able to work across technical and nontechnical teams, explain trade-offs, gain stakeholder buy-in, and operate with a growth mindset. They will work with another teammate and leverage previous work rather than start from scratch. If they are technically great, but not good with business trade offs & stakeholder communications, they aren't a good fit.

\n


\n

AI & OCR Solution Development

\n

    \n
  • Design, develop, and optimize OCR and Document Intelligence solutions using traditional OCR tools, cloud-native services, and Large Language Models (LLMs).
  • \n

  • Build and enhance document extraction pipelines for structured and unstructured documents, including handwritten forms, poor-quality faxes, scanned images, and complex business documents.
  • \n

  • Evaluate and compare OCR technologies such as Tesseract, cloud OCR services, and LLM-based extraction approaches to identify the best solution for business needs.
  • \n

  • Develop processes to improve extraction accuracy and data quality through validation, feedback loops, and continuous model optimization.
  • \n

\n

Data Quality & Validation

\n

    \n
  • Design automated validation frameworks to compare extracted data against databases and business rules.
  • \n

  • Analyze OCR failures and identify root causes related to image quality, handwriting recognition, document layouts, and extraction limitations.
  • \n

  • Develop strategies to improve data extraction accuracy and downstream business outcomes.
  • \n

\n

Application Development & Deployment

\n

    \n
  • Develop, deploy, and maintain scalable production applications and APIs supporting document processing workflows.
  • \n

  • Build end-to-end solutions using Python, SQL, and cloud technologies.
  • \n

  • Implement monitoring, logging, alerting, and performance optimization for production workloads.
  • \n

  • Collaborate with platform and engineering teams to ensure reliability, security, and scalability.
  • \n

\n

Technical Leadership

\n

    \n
  • Provide technical guidance on architecture decisions, tool selection, and implementation approaches.
  • \n

  • Assess emerging technologies in OCR, Generative AI, and Document Intelligence.
  • \n

  • Drive technical excellence through best practices, code reviews, testing, and documentation.
  • \n

  • Leverage existing solutions and prior work to accelerate delivery rather than rebuilding from scratch.
  • \n

\n

Business & Stakeholder Engagement

\n

    \n
  • Partner closely with business stakeholders to understand requirements, priorities, and success metrics.
  • \n

  • Translate business challenges into technical solutions that deliver measurable value.
  • \n

  • Clearly communicate technical trade-offs, risks, assumptions, and recommendations to both technical and non-technical audiences.
  • \n

  • Build consensus and gain stakeholder buy-in for proposed approaches.
  • \n

\n


\n

Qualifications

\n

    \n
  • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, Information Systems, or a related field.
  • \n

  • 5+ years of experience in AI/ML, OCR, Document Processing, Data Engineering, or related domains.
  • \n

  • Hands-on experience with OCR technologies including:
  • \n

  • Tesseract OCR
  • \n

  • Cloud OCR platforms
  • \n

  • LLM-based document extraction solutions
  • \n

  • Intelligent Document Processing (IDP) platforms
  • \n

  • Experience processing:
  • \n

  • Handwritten documents
  • \n

  • Low-quality scans
  • \n

  • Faxes
  • \n

  • Images with artifacts and distortions
  • \n

  • Large-scale document repositories
  • \n

  • Strong proficiency in Python and SQL.
  • \n

  • Experience with cloud platforms, preferably Google Cloud Platform (GCP).
  • \n

  • Experience developing, deploying, and supporting production-grade applications.
  • \n

  • Experience developing, deploying, and supporting production-grade applications.
  • \n

  • Strong understanding of data validation, data quality, and document extraction workflows.
  • \n

  • Experience integrating AI/ML solutions into business processes.
  • \n

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...