# Siddharth Vohra > Senior AI Research Engineer at AWS and M.S. Computer Vision candidate at Carnegie Mellon. Auditing where language and vision-language models fail when evidence is missing, misleading, or changed only in presentation. Site: https://siddvoh.com/. Last updated 9 October 2026. Siddharth Vohra, also known as Sidd Vohra, is a Senior AI Research Engineer at Amazon Web Services in Pittsburgh, Pennsylvania, and an M.S. Computer Vision candidate at Carnegie Mellon University (School of Computer Science, Robotics Institute). Hello, I'm Sidd! I'm Siddharth Vohra, a Senior AI Research Engineer based in Pittsburgh, PA. I build AI-native products at Amazon Web Services and am pursuing a Master's in Computer Vision at Carnegie Mellon University's School of Computer Science (Robotics Institute). Before this, I built production ML systems for AWS Bedrock, Amazon Transcribe, and Amazon Translate, working across speech and language. I earned my Bachelor's in Mathematics-Computer Science from UC San Diego in 2022. ## Contact and profiles - Email: siddvoh@gmail.com - Google Scholar: https://scholar.google.com/citations?user=N9DDnEEAAAAJ - ORCID: https://orcid.org/0009-0002-6199-0485 - ResearchGate: https://www.researchgate.net/profile/Siddharth-Vohra-2 - GitHub: https://github.com/siddvoh - LinkedIn: https://www.linkedin.com/in/siddvoh - OpenReview: https://openreview.net/profile?id=~Siddharth_Vohra1 - Semantic Scholar: https://www.semanticscholar.org/author/Siddharth-Vohra/2057336758 - DBLP: https://dblp.org/pid/272/9316.html - Hugging Face: https://huggingface.co/siddvoh - CMU Robotics Institute: https://www.ri.cmu.edu/ri-people/siddharth-vohra/ - X: https://x.com/siddvoh - ORCID iD: 0009-0002-6199-0485 ## Press ### CMU Robotics Institute, 20 July 2026 "Healthcare Blind Spots: AI Models Prone To Fabricating Diagnoses" - Link: https://www.ri.cmu.edu/healthcare-blind-spots-ai-models-prone-to-fabricating-diagnoses/ Robotics Institute feature on the Hearsay study: asked to describe a medical image that was never attached, Claude, GPT-5 and Gemini fabricated a diagnosis in 18% of about 11,700 responses, shaped by the patient's stated age, gender and race. ### The National, 27 July 2026 "Patients warned off using AI chatbots for self-diagnosis as flaws revealed" - Link: https://www.thenationalnews.com/news/uae/2026/07/27/patients-warned-off-using-ai-chatbots-for-self-diagnosis-as-flaws-revealed/ Report on the same study: AI chatbots fabricated a diagnosis in 18 per cent of cases when the medical image was left out, and a diagnosis can change when only race, age or gender is written into the prompt. ### The Business Journals, 30 September 2026 "Companies grapple with how to manage workers and AI errors" - Link: https://www.bizjournals.com/bizjournals/news/2026/09/30/ai-tools-tips-managers.html Workplace feature that opens with the study and quotes Vohra on fixing the structure around an AI tool when it makes a mistake. ## Experience ### Amazon Web Services (Aug 2022 — Present) #### Senior AI Research Engineer - Period: Oct 2026 — Present - Location: Pittsburgh, PA - Team: AWS Frontier AI Engineering and Services #### AI Research Engineer - Period: Sep 2025 — Sep 2026 - Location: Pittsburgh, PA - Team: AWS Frontier AI Engineering and Services - Title history: Machine Learning Engineer until Aug 2026 Sole engineer on an AI-native engineering platform and multi-agent system that lets enterprise and public-sector teams create, deploy, govern and audit AI agents in one place. Lead agentic AI projects for AWS partners and customers, running working sessions with C-suite executives at a European insurer and a US public-sector energy organization. #### Software Development Engineer II - Period: Apr 2024 — Sep 2025 - Location: Seattle, WA - Team: AWS AI · Bedrock Generative AI Services Feature lead across Amazon Transcribe, Bedrock Data Automation and Amazon Translate. Built speech-recognition models up to 30% faster and 2x more accurate on multilingual transcription, and served as technical lead for Amazon Translate across 30+ enterprise customers. #### Software Development Engineer - Period: Aug 2022 — Apr 2024 - Location: Seattle, WA - Team: Lambda · App Runner · Elastic Beanstalk - Related: Using WAF with App Runner in Copilot, AWS Copilot CLI Blog: https://aws.github.io/copilot-cli/blogs/apprunner-waf/ Led the AWS WAF integration into the Copilot CLI so App Runner applications ship secure by default, built an automated testing suite for the App Runner console, and helped launch App Runner in three new regions. ### Teradata (Jul 2021 — Sep 2021) #### Software Engineer Intern - Period: Jul 2021 — Sep 2021 - Location: San Diego, CA Restructured large-object storage in the TeraCloud architecture, improving computation efficiency by roughly 50%. ## Publications Each entry states the paper's standing. Accepted papers are marked as not yet published, and non-archival work is cited by its arXiv preprint. ### Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines - Authors: Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu - Venue: Grounding Language Models: Learning Faithfully and Efficiently (GroundLM) Workshop at EMNLP 2026 - Status: Accepted, not yet published. Accepted to the GroundLM workshop at EMNLP 2026, to be presented on 29 October 2026 in Budapest. Until the proceedings appear, the arXiv preprint is the version to cite. - Record dated: 4 September 2026 - Page on this site: https://siddvoh.com/identity-handoffs/ (markdown: https://siddvoh.com/identity-handoffs/index.md) - Record to cite: https://arxiv.org/abs/2609.04579 - arXiv: https://arxiv.org/abs/2609.04579 - OpenReview: https://openreview.net/forum?id=hSlA52KOs0 - PDF: https://arxiv.org/pdf/2609.04579 - Code and data: https://github.com/siddvoh/rop - BibTeX: https://siddvoh.com/bib/rop.bib Traces whether an item selected early in a grounded QA pipeline survives retrieval and reaches the answer model. On 1,463 aligned HybridQA records, body-only BM25 dropped it out of the top five 26.6% of the time; hybrid reranking cut that to 1.0%. Abstract: Grounded language-model pipelines can be divided into three stages: selecting an object, retrieving passages for it, and using that evidence to answer. If the selected object must reach the reader, losing it breaks the handoff. Benchmark recall checks the dataset-linked object, which can differ. We audit 600 HybridQA questions across three selector families. On 1,463 resolvable records where the selected object matches the dataset-traced passage, exact key lookup and exact title matching return the object every time. With every ranked rule given the same decoded selected title, body-only BM25 omits it on 389 records (26.6%) at cutoff five, while hybrid retrieval with reranking omits it on 14 (1.0%). The two identities differ on 329 of 1,792 resolvable records. With original-question rankings, their top-five checks disagree on 106 records (5.9%). Frozen reader comparisons associate the aligned object's presence with 28.6 to 31.0 points higher exact match. In a deliberately selected 64-item cohort, removing that passage sharply lowers exact match, while removing a similar-length comparison passage does not reproduce the drop. We release the Returned-Object Profile (ROP), an executable record of the target, returned-ID field, cutoff, membership rule, and complete expected population, with data and an offline replay. ### Research Agents Feed Doubts, Not Beliefs: Auditing Belief Framing in Live-Web Search - Authors: Siddharth Vohra, Min Xu - Venue: Second Workshop for Research on Agent Language Models (REALM) at EMNLP 2026 - Status: Accepted, not yet published. Accepted as a Spotlight paper at the REALM workshop at EMNLP 2026, to be presented on 29 October 2026 in Budapest. - Selection: Spotlight (https://realm-workshop.github.io/accepted_papers/) - Record dated: 13 September 2026 - Page on this site: https://siddvoh.com/framedfacts/ (markdown: https://siddvoh.com/framedfacts/index.md) - Record to cite: https://siddvoh.com/framedfacts/ - OpenReview: https://openreview.net/forum?id=rto8fWpik1 - PDF: https://siddvoh.com/framedfacts/framedfacts.pdf - Code and data: https://github.com/siddvoh/framedfacts - BibTeX: https://siddvoh.com/bib/framedfacts.bib Four web-search agents tested on 48 expert-checked health claims across 1,874 runs. Stating doubt shifted every system toward rejection and more than doubled searches seeking disconfirming evidence. Replay tests traced roughly three quarters of the shift to how agents read the evidence they retrieved. Abstract: People increasingly hand fact-finding to LLM agents that search the web and return a cited report. We ask whether the belief a user states when making the request changes the research itself. We give 48 expert-verified health claims to four search-equipped agents under believer, skeptic, and neutral framings, analyzing 1,874 of 2,064 registered runs plus second-round arms reported separately. The effect is asymmetric: stating belief leaves conclusions unchanged (0.02, claim-clustered 95% CI −0.04 to 0.09), while stating doubt moves them toward rejection (−0.26, CI −0.39 to −0.15). Across opposing framings the same agent's conclusion shifts 0.26 points on a five-point scale (CI 0.15 to 0.40), a quarter of one label step, and every system moves the same way. The shift sits on contested claims (0.61) and is small on the settled tiers (0.10 and 0.11). Strictly opposite conclusions were rare and their registered paired test was null (p = 0.250), so this is a shift in degree rather than a reversal of verdicts. The trajectories mirror it: skeptic framing more than doubles deny-seeking queries (net −9.9 points against +0.5 for believer framing). Crossing the retrieved evidence between framings splits the gap into how the agent reads what comes back (0.23, CI 0.12 to 0.34) and which evidence comes back at all (0.08, CI 0.00 to 0.16), two channels that add to the end-to-end gap by construction and put about three quarters of it on reading, though the evidence interval reaches zero. Labeling retrieved results against the claim finds no evidence that stated doubt returns a more hostile pool (1.7 points, CI −1.5 to 5.5). An indicative post-hoc follow-up on the 32 claims it reached adds a both-sides instruction to the skeptic condition, where the bias lives, and shifts conclusions by 0.16 (CI 0.04 to 0.30). We release FramedFacts-48: claims, prompts, and analyzed runs. ### The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits - Authors: Siddharth Vohra, Manikandan Ravikiran - Venue: Second Workshop for Research on Agent Language Models (REALM) at EMNLP 2026 - Status: Accepted, not yet published. Accepted to the REALM workshop at EMNLP 2026, to be presented on 29 October 2026 in Budapest. Until the proceedings appear, the arXiv preprint is the version to cite. - Record dated: 8 September 2026 - Page on this site: https://siddvoh.com/biasinaudits/ (markdown: https://siddvoh.com/biasinaudits/index.md) - Record to cite: https://arxiv.org/abs/2609.09048 - arXiv: https://arxiv.org/abs/2609.09048 - OpenReview: https://openreview.net/forum?id=nYzDm5aD60 - PDF: https://arxiv.org/pdf/2609.09048 - BibTeX: https://siddvoh.com/bib/biasinaudits.bib Across 40,726 requests to five models in hiring, lending and medical triage, an earlier finding that models favour minority applicants when rating and penalise some when ranking did not appear, and none of 36 planned comparisons held up after correction. Where an application sat in the list moved the rankings as much as any name difference measured. Abstract: Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias. ### When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA - Authors: Manikandan Ravikiran, Siddharth Vohra - Venue: Grounding Language Models: Learning Faithfully and Efficiently (GroundLM) Workshop at EMNLP 2026 - Status: Accepted, not yet published. Accepted to the GroundLM workshop at EMNLP 2026, to be presented on 29 October 2026 in Budapest. Until the proceedings appear, the arXiv preprint is the version to cite. - Record dated: 8 September 2026 - Record to cite: https://arxiv.org/abs/2609.08934 - arXiv: https://arxiv.org/abs/2609.08934 - OpenReview: https://openreview.net/forum?id=2tSkWhDBrE - PDF: https://arxiv.org/pdf/2609.08934 - BibTeX: https://siddvoh.com/bib/defer.bib Across 220,000 responses from four models in English, Hindi, Bengali, Tamil and Telugu, an unverified expert cue pulled models to a named wrong option in 41.1% of trials they had first answered correctly. A majority cue did so in 12.5%. Abstract: Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce neutral-conditioned misleading cue adoption rate (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220,000 outputs, the expert template yields 41.1% aggregate NC-MCAR, compared with 12.5% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence. ### Absent-Byte Diagnoses: Auditing Structured Medical VLM Interfaces - Authors: Siddharth Vohra, Manikandan Ravikiran - Venue: MedAgent: 2nd Agentic AI for Medicine Workshop at MICCAI 2026 - Status: Published. Open-access version published by the MICCAI Society in the MICCAI 2026 satellite event proceedings (MedAgent, the 2nd Agentic AI for Medicine workshop). The Springer LNCS volume is forthcoming. - Record dated: 21 September 2026 - Page on this site: https://siddvoh.com/absent-byte/ (markdown: https://siddvoh.com/absent-byte/index.md) - Record to cite: https://papers.miccai.org/miccai-2026-sat/MedAgent_050.html - OpenReview: https://openreview.net/forum?id=fIfFrQzvry - PDF: https://papers.miccai.org/miccai-2026-sat/paper/MedAgent_050.pdf - Code and data: https://github.com/siddvoh/absent_byte - BibTeX: https://siddvoh.com/bib/absentbyte.bib Three of five medical vision-language models returned a structured diagnosis in 617 of 3,000 calls where the prompt claimed an image was attached but the request carried no image bytes. A client-side verifier blocked all 5,047 tested evidence-binding violations. Abstract: When a medical agent calls a vision-language model, prompt text, image bytes, and structured response fields may pass through separate software components. We study a mismatch at this boundary. The prompt says an image is attached, while the retained client request contains no image bytes. A diagnosis in this state can pass a schema check and reach later software without visual evidence. In the baseline contract, GPT-5.4, Claude, and Gemini filled the diagnosis field in 617 of 3,000 calls across five hosted models. Changing only demographic wording also changed the returned diagnosis distribution. In one matched GPT-5.4 chest X-ray pair, the leading diagnosis moved from Pneumothorax in 54 of 100 calls for a profile described as a white man to Sarcoidosis in 77 of 100 after white changed to Black. The broader analysis covers 288 direct age, race-word, and sex-word contrasts. None of 19,350 hosted no-attachment controls filled the diagnosis field. Four public software regressions show how images or task text can disappear at client boundaries. We bind the completed request and selected diagnosis field to a caller-owned record of the task and ordered image fingerprints. Three implementations pass all 1,118 controls and block all 5,047 specified violations in a 6,165-case offline suite. ### Hearsay: Vision-Language Medical Diagnoses Without an Image - Authors: Siddharth Vohra - Venue: 1st Workshop on Toward Trustworthy Vision-Language Models in the Wild (TrustVLM) at ACM ICMR 2026 - Status: Presented, non-archival. Peer-reviewed and presented at the TrustVLM workshop at ACM ICMR 2026 in Amsterdam. The workshop is non-archival, so the arXiv preprint is the version to cite. - Record dated: 29 July 2026 - Page on this site: https://siddvoh.com/hearsay/ (markdown: https://siddvoh.com/hearsay/index.md) - Record to cite: https://arxiv.org/abs/2607.26886 - arXiv: https://arxiv.org/abs/2607.26886 - OpenReview: https://openreview.net/forum?id=5LcHzZPUzS - PDF: https://arxiv.org/pdf/2607.26886 - BibTeX: https://siddvoh.com/bib/hearsay.bib Audits frontier vision-language models under missing-image medical prompts and identifies structured diagnostic confabulation, demographic sensitivity, and structured-output failure modes. Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern." GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension. ### TEEMIL: Towards Educational MCQ Difficulty Estimation in Indic Languages - Authors: Manikandan Ravikiran, Siddharth Vohra, Rajat Verma, Rohit Saluja, Arnav Bhavsar - Venue: 31st International Conference on Computational Linguistics (COLING 2025) - Status: Published. Published in the main conference proceedings of COLING 2025, the 31st International Conference on Computational Linguistics. - Record dated: January 2025 - Record to cite: https://aclanthology.org/2025.coling-main.142/ - PDF: https://aclanthology.org/2025.coling-main.142.pdf - BibTeX: https://siddvoh.com/bib/teemil.bib Introduces the TEEMIL-H and TEEMIL-K datasets for Hindi and Kannada MCQ difficulty estimation, with multilingual model baselines and ablations over context, answer options and none-of-the-above options. Abstract: Difficulty estimation of multiple-choice questions (MCQs) is crucial for creating effective educational assessments, yet remains underexplored in Indic languages like Hindi and Kannada due to the lack of comprehensive datasets. This paper addresses this gap by introducing two datasets, TEEMIL-H and TEEMIL-K, containing 4689 and 4215 MCQs, respectively, with manually annotated difficulty labels. We benchmark these datasets using state-of-the-art multilingual models and conduct ablation studies to analyze the effect of context, the impact of options, and the presence of the None of the Above (NOTA) option on difficulty estimation. Our findings establish baselines for difficulty estimation in Hindi and Kannada, offering valuable insights into improving model performance and guiding future research in MCQ difficulty estimation. ### You Reap What You Sow—Revisiting Intra-class Variations and Seed Selection in Temporal Ensembling for Image Classification - Authors: Manikandan Ravikiran, Siddharth Vohra, Yuichi Nonaka, Sharath Kumar, Shibashish Sen, Nestor Mariyasagayam, Kingshuk Banerjee - Venue: International Conference on Frontiers in Computing and Systems (COMSYS 2021) - Status: Published. Published by Springer in the COMSYS 2021 proceedings (Lecture Notes in Networks and Systems, vol. 404, pages 73 to 82). First online 28 June 2022; the volume is dated 2023. - Record dated: 28 June 2022 - Record to cite: https://doi.org/10.1007/978-981-19-0105-8_8 - DOI: https://doi.org/10.1007/978-981-19-0105-8_8 - BibTeX: https://siddvoh.com/bib/reap.bib Studies how intra-class variability, seed size and seed selection affect semi-supervised Temporal Ensembling performance across image-classification datasets. Ravikiran and Vohra contributed equally; names are ordered alphabetically. Abstract: In this work, we present our study on the influence of intra-class variations and seed selection on image classification using temporal ensembling. Through our experiments, we observe that for Fashion-MNIST and Kuzushiji-MNIST datasets with medium and high intra-class variations, (a) classification accuracy declines by 20% and 30%, (b) raising seed images by 5x improves accuracy by 15% and 8% respectively. Additionally, preliminary investigation on the Fashion-MNIST dataset reveals that with diversified seed selection class, level accuracy improves by 6%. In due process, our study exhibits convergence issues for unsupervised loss when training on datasets with medium and high intra-class variations. ### Investigating the Effect of Intraclass Variability in Temporal Ensembling - Authors: Siddharth Vohra, Manikandan Ravikiran - Venue: arXiv - Status: Preprint. arXiv preprint, posted 20 August 2020. It led to the Springer paper published in 2023. - Record dated: 20 August 2020 - Record to cite: https://arxiv.org/abs/2008.08956 - arXiv: https://arxiv.org/abs/2008.08956 - PDF: https://arxiv.org/pdf/2008.08956 - BibTeX: https://siddvoh.com/bib/intraclass.bib Early study of how within-class variation and the number and choice of labelled seeds affect Temporal Ensembling. Led to the 2023 Springer paper above. Abstract: Temporal Ensembling is a semi-supervised approach that allows training deep neural network models with a small number of labeled images. In this paper, we present our preliminary study on the effect of intraclass variability on temporal ensembling, with a focus on seed size and seed type, respectively. Through our experiments we find that (a) there is a significant drop in accuracy with datasets that offer high intraclass variability, (b) more seed images offer consistently higher accuracy across the datasets, and (c) seed type indeed has an impact on the overall efficiency, where it produces a spectrum of accuracy both lower and higher. Additionally, based on our experiments, we also find KMNIST to be a competitive baseline for temporal ensembling. ## Academic service ### Conference peer review Reviewer and programme committee member for workshops and affinity events at NeurIPS, ICML, EMNLP, ECCV, MICCAI, COLM, IJCAI-ECAI and ACM KDD, and ethics reviewer for the NeurIPS 2026 main conference, covering mechanistic interpretability, biomedical retrieval, and the robustness and safety of agentic AI systems. - NeurIPS 2026: Ethics reviewer, main conference. Workshops: Symmetry and Geometry in Neural Representations; GlobalSouthAI (affinity workshop); Trustworthy AI for Good; AI for Meta-Science: Scaling and Organizing Science in the Age of AI Scientists; AI for Drug Discovery: Bridging the Translation Gap; Interpretability for Discovery: Understanding and Discovering Novel Knowledge in AI Models; Mathematical Reasoning and AI; Grounded and Faithful Vision-Language Models for Real-World Deployment; Continual Learning for Enterprise AI Agents; Dynamics at the Frontiers of Optimization, Sampling, and Games; AI and the Self: Human Identity, Authenticity and Agency in the Age of AI - EMNLP 2026: Workshops: Grounding Language Models: Learning Faithfully and Efficiently; Workshop for Research on Agent Language Models - COLM 2026: Workshops: Context Beyond the Window: Persistent Knowledge in Language Models; Social Simulation with LLMs: Fidelity in Applications; Workshop on Efficient Reasoning - ECCV 2026: Workshops: Women in Computer Vision; Computer Vision for Ecology - ICML 2026: Workshop: Compositional Learning: Safety, Interpretability and Agents - MICCAI 2026: Workshop: Workshop on AI for Safe Surgery - ACM KDD 2026: Programme Committee. Workshop: Agentic Software Engineering (SE 3.0): The Rise of AI Teammates - IJCAI-ECAI 2026: Diversity and inclusion activity: Global South AI ### Hackathon judging Invited to the judging panels of flagship university hackathons at Carnegie Mellon, UC San Diego and Cincinnati. HackCMU and DiamondHacks each had about 100 submitted projects. - HackCMU 2025, ACM@CMU, Carnegie Mellon University, September 2025. 600+ participants, 100+ submissions. - DiamondHacks 2025, ACM at UC San Diego, April 2025. 650+ registrations, 95 projects. - RevolutionUC 2025, ACM@UC, University of Cincinnati, March 2025. 873 applications, 300 selected. - Hackathon Raptors: Elected member, Guild of Expert Engineers, August 2026. Elected by peer review. Election requires at least four of five votes from randomly selected sitting Fellows against eight published criteria. https://www.raptors.dev/fellow-membership ## Honors and memberships ### Gemini Academic Program Award, Google - Kind: Award or selection - When: 2026 $20,000 in Google Cloud credits for Gemini foundation-model research and multimodal evaluation audits. ### Member, Pittsburgh Section, IEEE - Kind: Membership - When: Membership for 2026 Member in good standing of the Institute of Electrical and Electronics Engineers, valid through December 2026. ### Professional Member, ACM - Kind: Membership - When: Member since May 2026 Admitted to the Association for Computing Machinery having fulfilled the requirements for Professional Membership. ## Education ### M.S. Computer Vision, Carnegie Mellon University (Expected Dec 2026) School of Computer Science, Robotics Institute. ### B.S. Mathematics-Computer Science, University of California, San Diego (2019 — 2022) Cum laude, Provost Honors all quarters. Founding and principal member of the Machine Learning Club. ## Reuse and citation Original text on this site is licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Papers, abstracts, PDFs and quoted headlines keep their own terms. If this work is useful to you, you are welcome to cite it. BibTeX for every paper is at https://siddvoh.com/papers.bib.