MOHAMMED.

Research / 02

Preprints, papers,
and ongoing work.

5 papers spanning knowledge-guided reinforcement learning, controllable protein design, interpretable clinical NLP, mechanistic safety of tool-using LLMs, and few-shot fault diagnosis under data scarcity. Click any thumbnail or title to open the full PDF.

01 / Themes

/01

RL & Knowledge

When structured priors help agents, and when they hurt.

/02

Protein Design

Controllable biophysics in inverse folding and binder design.

/03

Healthcare AI

Interpretable clinical NLP, concept-grounded diagnosis.

/04

LLM Safety

Channel-specific vulnerability, mechanistic interpretability.

/05

Applied Generative ML

Few-shot diagnosis, augmentation under data scarcity.

02 / Papers (5)
The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning preview
Submitted2026

Reinforcement Learning · Knowledge Graphs

The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning

Mohammed Sameer SyedAAAI 2027 · arXiv:2607.19616University of Arizona

Knowledge graphs (KGs) are widely used to inject prior knowledge into reinforcement learning, yet the literature is dominated by single-domain, positive-result method papers, so we lack a systematic account of when KG structure helps an agent, when it is neutral, and when it hurts. We run a controlled study that independently varies the RL task, the injection mechanism (state features, action masking, or potential-based reward shaping), and KG quality over a synthetic, fully controllable KG on MiniGrid. Structured guidance improves sample efficiency and solve reliability on compositional sparse-reward tasks, and a shuffle control that permutes the KG's edges while preserving their count collapses the benefit toward baseline (masking p=0.0001; shaping p=0.006), so the gain is structural rather than generic regularization. Most consequentially, safety depends on the mechanism: soft, optimality-preserving injection benefits from correct knowledge and harmlessly ignores incorrect knowledge, whereas hard masking is brittle and can make a wrong KG worse than no KG. A UMLS-derived clinical case study on MIMIC-IV sepsis management under offline RL is a careful null, underscoring that benefits require task structure the chosen mechanism can exploit.

70% → 97%

Solve reliability

seeds solved

p = 0.0001

Shuffle control

masking, d=1.08

6 × 2

Envs × learners

Knowledge GraphsReward ShapingAction MaskingMiniGridOffline RL
Controllable Electrostatic Protein Design: Dialing Net Charge at Inference Time in ProteinMPNN and BindCraft preview
Working Draft2026

Computational Biology · Controllable Generation

Controllable Electrostatic Protein Design: Dialing Net Charge at Inference Time in ProteinMPNN and BindCraft

Mohammed Sameer SyedTargeting PSB 2027University of Arizona

Deep inverse-folding models such as ProteinMPNN design sequences that fold to a target backbone, but expose no control over the electrostatic properties of their output: a user cannot request a sequence with a specified net charge. ProteinMPNN already exposes a per-residue logit bias; the missing piece is a way to aim it. We supply a closed-loop, per-protein calibration loop that turns the raw charged-residue bias into a controller hitting a specified net charge, a characterization of the charge-versus-foldability trade-off, and the transfer of both into a de-novo binder-design pipeline. On leakage-free, sequence-clustered held-out data the calibrated controller hits charge targets with a mean absolute error of 7–10 charge units, where a single global bias leaves a protein-dependent residual of tens of units, and folding is preserved within a well-characterized band. Integrating the controller into BindCraft, we design binders against PD-L1 at a specified charge: a gentle setting shifts binder net charge from −6.0 to ~0 with no significant drop in interface confidence, whereas an over-aggressive setting degenerates the binder into poly-lysine and collapses binding. Charge thus becomes a tunable design dial with a characterized usable band.

7–10

Charge MAE

held-out, leakage-free

−6.0 → ~0

PD-L1 binder charge

ipTM 0.42 vs 0.46

1.69 Å

Median scRMSD

fold preserved in band

ProteinMPNNBindCraftInverse FoldingNet ChargeESMFold
ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding preview
Preprint2025

Healthcare AI · Interpretability

ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding

Mohammed Sameer Syed, Xuan LuUniversity of Arizona

Automated ICD-10 coding from clinical discharge summaries requires models that are both accurate on long-tailed multi-label classification tasks and interpretable to clinicians. We present ShifaMind, a concept-grounded architecture built around a Multiplicative Concept Bottleneck (MCB), which changes the form, rather than the width, of the bottleneck. Instead of projecting through a narrow concept layer, ShifaMind uses a learned multiplicative gate over a concept-grounded representation while retaining a scalar concept interface for inspection. On MIMIC-IV top-50 ICD-10 coding, ShifaMind achieves performance competitive with the strongest baseline LAAT across F1, AUC, and ranking metrics, while outperforming five additional ICD-coding baselines and providing concept-mediated explanations.

0.712

Macro-F1

MIMIC-IV top-50

4.3×

over Vanilla CBM

0.704

CSTPR

Concept BottleneckClinical NLPICD-10MIMIC-IVInterpretability
Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models preview
Preprint2025

LLM Safety · Mechanistic Interpretability

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

Mohammed Sameer Syed, Rozhin YasaeiUniversity of Arizona

As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party content, their attack surface expands well beyond what users type. We introduce the Safety Asymmetry Score (SAS), which measures how much a model's susceptibility to adversarial content shifts depending on whether that content arrives in the user message, tool metadata, or tool output, using matched payload pairs that keep the malicious text identical and vary only the context of delivery. Evaluated across 6 production LLMs and three attack families, agent-native models are substantially more vulnerable when adversarial content arrives via tool descriptions than via user messages, while general-purpose models show the reverse. A mechanistic study on Llama 3.3 70B reveals that the safety-relevant representation is causally present at mid-to-late network depths but non-linearly encoded, explaining why linear probes fail to detect it.

+30.4 pp

Group SAS gap

6 / 98

Models · cases

ρ = 0.54

vs MCPTox

LLM SafetyTool UseMCPActivation PatchingLlama 3.3
Fault Diagnosis of Power Transformer Using Frequency Response Analysis with the aid of SpectralGAN-Augmented Transformer Neural Network preview
Submission Pending2025

Power Systems · Generative Models

Fault Diagnosis of Power Transformer Using Frequency Response Analysis with the aid of SpectralGAN-Augmented Transformer Neural Network

Mohammed Sohail Syed, Madhava Rao Trilingi, Mohammed Sameer SyedIEEE Transactions on Power DeliverySRM University · University of Arizona

Accurate classification of power transformer winding deformation faults from frequency response analysis (FRA) measurements is constrained by the fundamental scarcity of labelled fault data. Existing data-driven methods either overfit to small training sets or rely on evaluation protocols that expose the test sample to the generative model, producing optimistic accuracy estimates. This paper presents a three-stage diagnostic pipeline that directly addresses both limitations. First, 48-dimensional indicator vectors are extracted from three IEC-standard sub-bands of the measured FRA transfer function. Second, SpectralGAN—a conditional WGAN-GP whose critic employs spectral normalisation on every linear layer—synthesises class-conditional indicator vectors from as few as 20 training samples per fold. Third, FRATransformer, a lightweight multi-head self-attention classifier that treats each sub-band feature block as a distinct token, classifies a mixed corpus of simulated samples, Gaussian jitter copies, and GAN-generated data. Evaluated under strict 21-fold Leave-One-Out Cross-Validation on a 21-sample simulated dataset spanning healthy, axial displacement, and radial deformation classes—with no test sample participating in GAN or classifier training—the pipeline achieves 85.7% accuracy and macro F1 = 0.838, a +23.8 pp improvement over the best SVM baseline (61.9%), with a 95% Clopper-Pearson confidence interval of [63.7%, 97.0%]. A four-condition ablation reveals that Gaussian jitter augmentation is the dominant driver of performance gain, while SpectralGAN synthesis further alters error patterns without changing total error count at this dataset scale.

85.7%

Accuracy

21-fold LOOCV · 95% CI [63.7, 97.0]

+23.8 pp

over SVM baseline

0.838

Macro F1

WGAN-GPSpectral NormalisationSelf-AttentionFRAFew-Shot