Method Article

Identifying Immune-related Molecular Biomarkers in Autism Spectrum Disorder Using Data-independent Acquisition Proteomics and Machine Learning

DOI:

10.3791/68949

September 26th, 2025

* These authors contributed equally

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Here, we present a protocol using data-independent acquisition mass spectrometry and machine learning that identified eight immune-related proteins as accurate biomarkers for early Autism spectrum disorder diagnosis, validated by an enzyme-linked immunosorbent assay.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study presents a reproducible protocol for identifying serum protein biomarkers associated with autism spectrum disorder (ASD) using data-independent acquisition (DIA) mass spectrometry combined with machine learning (ML). DIA enables unbiased, high-resolution profiling of the serum proteome, including low-abundance proteins, while ensuring reproducibility across samples. ML approaches were applied to select diagnostically informative protein panels and improve model robustness. The analysis included serum from 99 children with ASD and 70 age-matched controls. High-abundance proteins were depleted, peptides were prepared using standardized digestion and fractionation procedures, and DIA was performed on a high-resolution mass spectrometer. Data processing and quantification identified differentially expressed proteins, which underwent functional enrichment analysis. Eight immune-related proteins emerged as strong candidates for biomarker development. A logistic regression model trained on these proteins achieved 95.27% accuracy, a Kappa value of 0.9025, and an AUC of 1.000 in cross-validation. These findings demonstrate the potential of DIA-based proteomics, combined with machine learning, as a robust framework for biomarker discovery in ASD and for adaptation in broader clinical research.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Autism spectrum disorder (ASD) is a group of early-onset neurodevelopmental disorders characterized by heterogeneity in etiology and clinical presentation. Core features include persistent deficits in social communication and interaction, as well as restricted, repetitive behaviors, interests, or activities. In the United States, prevalence is approximately 2.3% among 8-year-old children and around 2.2% among adults, underscoring its public health impact1,2,3,4. Risk factors are diverse, including genetic predispositions, immune dysregulation, and prenatal environmental exposures5,6,7. Early diagnosis and intervention can significantly improve developmental outcomes, making the identification of objective and reliable biomarkers a major focus of ASD research8,9,10. This protocol builds on our previously published work applying data-independent acquisition (DIA) proteomics and machine learning to identify immune-related proteins as potential biomarkers for early ASD diagnosis11.

Despite extensive efforts, no specific and universally validated biomarkers currently exist for clinical ASD diagnosis12. Proposed candidates-such as alterations in the gut microbiome13, elevated interleukin-6 (IL-6)14, changes in brain-derived neurotrophic factor (BDNF)15, and oxidative stress markers like glutathione16-remain preliminary and lack reproducibility for clinical use. Proteomics has emerged as a promising approach for identifying disease-specific molecular signatures, and several studies have investigated different biological samples (blood, saliva, urine, PBMCs) for differentially expressed proteins8,17,18,19,20,21,22. For instance, Bao et al. demonstrated that inflammatory proteins identified by Olink proteomics may aid in early ASD diagnosis (17), while other studies suggest that shared proteomic and metabolic pathways may yield robust biomarkers despite ASD's genetic heterogeneity23.

DIA mass spectrometry has gained increasing attention for its comprehensive and reproducible protein profiling. Unlike traditional data-dependent acquisition (DDA), which selectively fragments the most intense ions, DIA fragments all precursor ions across predefined m/z windows. This provides deeper proteome coverage and improved reproducibility across large cohorts, a key advantage for clinical comparisons14. Benchmarking studies show that DIA detects more quantifiable peptides than DDA, particularly for low-abundance proteins, with lower inter-run variation14.

Building on these advances, we applied DIA-based proteomic analysis to serum samples from 99 children with ASD and 70 controls, following depletion of high-abundance proteins. Our findings highlight the potential of immune-related proteins as molecular markers for early ASD diagnosis and demonstrate the value of DIA-based proteomics in biomarker discovery when combined with rigorous methodology11.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The protocol was conducted in accordance with the Declaration of Helsinki and the protocol was approved by the Institutional Review Board at Changsha Maternal and Child Health Hospital; informed consent was obtained from the subjects.

1. Identification of children with autism with DSM-5

  1. Collecting medical history and background information
    1. Developmental history
      1. Collect information on the patient's early development, including language, social, and motor skill progression.
      2. Note any developmental delays or abnormalities (e.g., language delay, difficulties in social interaction).
    2. Family history
      1. Inquire about any family history of autism or other neurodevelopmental disorders.
    3. Current functional level
      1. Assess the patient's performance in daily life, including learning, work, social interactions, and independent living skills.
  2. Using DSM-5 diagnostic criteria
    1. Persistent deficits in social communication and social interaction
      1. Ensure that at least two of the following three criteria are met:
        1. Deficits in social-emotional reciprocity-look for a lack of normal eye contact, facial expressions, or body language and difficulty forming age-appropriate friendships or relationships.
        2. Deficits in nonverbal communicative behaviors-look for challenges in using gestures, facial expressions, or tone of voice to convey emotions and a limited understanding of nonverbal cues from others.
        3. Deficits in developing, maintaining, and understanding relationships-look for difficulty adapting to different social contexts and a lack of interest in peers or inability to engage in imaginative play.
    2. Restricted, repetitive patterns of behavior, interests, or activities
      1. Ensure that at least two of the following four criteria are met:
        1. Stereotyped or repetitive motor movements (e.g., hand flapping, body rocking, or repetitive object use).
        2. Insistence on sameness or ritualized patterns of behavior -look for extreme distress over minor changes in routine.
        3. Highly restricted, fixated interests-look for abnormally intense focus on specific topics or activities.
        4. Hyper- or hyporeactivity to sensory input-look for atypical responses to sensory stimuli such as sounds, lights, or touch.
  3. Assessment of symptom onset and severity
    1. Timing of symptoms-confirm that symptoms were present in early childhood (typically before age 3), even if they become more apparent later.
    2. Impact of symptoms-confirm that the symptoms cause significant impairment in social, occupational, or other important areas of functioning.
    3. Severity levels
      NOTE: According to DSM-5, ASD severity is categorized into three levels (Supplemental Table S1).
      1. Categorize as level 1 if the patient only requires mild support.
      2. Categorize as level 2 if the patient requires substantial support (moderate).
      3. Categorize as level 3 if the patient requires very substantial support (severe).
  4. Exclusion of other potential causes
    1. Medical Examination Conduct necessary medical evaluations (e.g., genetic testing, brain imaging) to rule out other conditions that may cause similar symptoms (e.g., genetic syndromes, hearing impairments, intellectual disabilities).
    2. Comorbidity assessment: Evaluate the presence of comorbid conditions (e.g., attention deficit hyperactivity disorder, anxiety disorders, depression, epilepsy, etc.).

2. Sample preparation for DIA mass spectrometry analysis

  1. Ethical compliance and sample collection
    1. Obtain informed consent from the parents or legal guardians of children aged 3-7 years diagnosed with Autism Spectrum Disorder (ASD).
    2. Classify patients into severity levels 1 to 3 according to the diagnostic criteria outlined in the American DSM-5 for autism (step 1.3.3).
    3. Collect serum samples from the participants. Ensure that all samples are processed within four hours of blood collection to prevent protein degradation. Keep samples on ice during processing.
  2. Removal of high-abundance proteins
    1. Use a commercial kit to deplete high-abundance proteins from 60 µL of serum per sample, following the manufacturer's instructions. Briefly, equilibrate the depletion column with binding buffer, load the serum sample, and allow it to pass through the column under gravity flow. Collect the flowthrough, which contains the low-abundance protein fraction.
    2. Measure the total protein concentration using a BCA assay. Normalize all samples to a final concentration of 0.5-1.0 µg/µL prior to in-solution digestion. Ensure that each sample contains at least 100 µg of protein for subsequent analysis.
  3. Protein digestion
    NOTE: Protein digestion was performed by using the FASP method described by Wisniewski et al.24.
    1. Add the detergent, Dithiothreitol (DTT), and Iodoacetamide (IAA) in UA (Urea buffer) buffer to block reduced cysteine.
    2. Digest the protein suspension with trypsin at a 50:1 ratio overnight at 37 °C.
  4. Peptide Desalting, Cleanup, and High-pH Reversed-Phase Fractionation
    1. Centrifuge the peptide mixtures at 16,000 × g for 15 min at  °C to remove insoluble debris.
    2. Transfer the supernatant (containing digested peptides) to a new low-bind microcentrifuge tube to minimize adsorption losses.
    3. .Prepare C18 microcolumns (in-house packed with C18 resin) by preconditioning with 100% methanol (20 µL) and equilibrating with 0.1% (v/v) trifluoroacetic acid (TFA) in water (buffer A; 20 µL).
    4. Load the peptide sample onto the microcolumn. Wash the column with 20 µL of buffer A to remove salts, detergents, and non-peptidic contaminants.
    5. Elute purified peptides with 20 µL of 80% acetonitrile containing 0.1% TFA.
    6. Dry the eluted peptides under vacuum using a centrifugal vacuum concentrator. Store dried peptides at -8 °C until further use.
    7. Reconstitute dried peptides in 0.1% formic acid prior to LC-MS/MS analysis.
    8. 2.4.8.Quantify peptide concentration by measuring absorbance at 280 nm (OD280) using a spectrophotometer, accounting for contributions from tryptophan and tyrosine residues for accurate quantification.
      To fractionate peptide mixtures using high-pH reversed-phase HPLC, use a C18 column (3.5 µm, 2.1 x 150 mm) on an HPLC system with a flow rate of 0.3 mL/min, mobile phase A: 10 mM ammonium formate in water, pH 10 (adjusted with ammonium hydroxide), mobile phase B: 10 mM ammonium formate in 90% acetonitrile, pH 10. Perform a gradient elution to collect 60 fractions per sample over ~60 min.
    9. Combine every third fraction to reduce redundancy, resulting in 20 pooled fractions per sample. Dry each pooled fraction under vacuum for downstream analysis.
      NOTE: The resulting peptide fractions are now ready for nano-LC-MS/MS analysis.

3. Submission for DIA mass spectrometry analysis

  1. DIA mass spectrometry analysis
    1. Spike the Data-dependent Acquisition (DDA) Peptide from the HPRP fraction with iRT standard peptides and separate them using reverse-phase high-performance liquid chromatography (RP-HPLC) on a nano-HPLC system with a column (75 µm x 150 mm; 2 µm C18 beads, 120 Å) at a flow rate of 300 nL/min with mobile phase A: 0.1% formic acid in water, mobile phase B: 0.1% formic acid in 95% acetonitrile.
    2. Elute the peptides over 60 min with a linear gradient of buffer B set as follows: 0 - 2 min, linear gradient from 2% to 5% buffer B; 2 - 42 min, linear gradient from 5% to 20% buffer B; 42 - 50 min, linear gradient from 20% to 35% buffer B; 50 - 52 min, linear gradient from 35% to 90% buffer B; 52 - 60 min, buffer B maintained at 90%.
    3. Analyze the eluted peptides on the referenced mass spectrometer. Acquire MS data using a data-dependent top20 method dynamically choosing the most abundant precursor ions from the survey scan (350 - 1500 m/z) for HCD fragmentation.
    4. Run the instrument with peptide recognition mode enabled. Use a lock mass of 445.120025 Da as the internal standard for mass calibration. Acquire the full MS scans at a resolution of 70,000 at m/z 200, and 17,500 at m/z 200 for MS/MS scan. Set the maximum injection time to 50 ms for MS and 30 ms for MS/MS, normalized collision energy to 28, the isolation window to 1.6 Th, and the dynamic exclusion duration to 30 s.
  2. LC-MS/MS analysis for data-independent acquisition (DIA)
    1. Spike the peptides from each sample with iRT equally and separately.
    2. Perform LC-MS/MS on a quadrupole mass spectrometer coupled with a nano-HPLC system. Set the LC condition in the same way as for the DDA method above. Perform a survey scan from 400 to 1,200 m/z at resolution 60,000 with AGC target of 3E6 and 30 ms injection time. Acquire the DIA MS/MS scans at resolution 15,000 with a 20 m/z isolation window and with an AGC target of 1E6 and 50 ms injection time. Set the normalized collision energy to 30.
    3. Record the spectra of full MS and DIA scans in profile and centroid types, respectively.
  3. Sequence database searching
    1. Analyze the DDA MS data using the DIA software2.
    2. Search the MS data against the UniProtKB human database (186,532 total entries, downloaded on 10/2019), spiked with proteins consisting of 11 iRT peptide sequences.
    3. Select trypsin as the digestion enzyme. Define the maximal two missed cleavage sites and the mass tolerance of 4.5 ppm for precursor ions and 20 ppm for fragment ions for database search. Define carbamidomethylation of cysteines as a fixed modification and acetylation of protein N-terminal and oxidation of Methionine as variable modifications for database searching.
    4. Filter the database search results and export them with <1% false discovery rate (FDR) at peptide-spectrum-matched and protein levels, respectively.
  4. Perform raw data processing
    1. Analyze the DIA MS data were analyzed with the DIA software [34, 35] for a spectral library generation from the search results. Use default settings for the search and the dynamic iRT for retention time prediction. Ensure that the interference correction for MS/MS scan is enabled.
    2. Export the results with <1% FDR at the peptide level.

4. Differential protein analysis

  1. Carry out hypothesis testing using Student's t-test combined with fold change (FC) at http://www.omickits.com/open/tooldetail?id=70.
    1. Log in to the cloud platform and navigate to the Hypothesis Testing Analysis tool. Upload the preprocessed protein quantification data file (e.g., CSV or TXT format).
    2. In the parameter settings, select Student's t-test as the statistical method and the significance threshold at p-value < 0.05. Define the fold change threshold as FC > 1.5 or FC < 1/1.5. Click Run Analysis and wait for the results to be generated.
    3. Download the output file containing p-values, log2(FC), and significance status for each protein.
      NOTE: This dual-criteria approach balances statistical significance with biological relevance, ensuring robust identification of differentially expressed proteins (DEPs).

5. Signal path analysis

  1. Volcano plot visualization
    1. Navigate to the tool at http://www.omickits.com/open/tooldetail?id=63 and then, to the volcano plot tool page.
      1. Upload the DEP analysis result file from section 4.
      2. Configure visualization parameter: X-axis: log2(Fold Change) - indicates direction of change; Y-axis: -log10(p-value) - reflects statistical significance; Color coding: Red : significantly upregulated proteins (p < 0.05 and FC > 1.5); Blue : significantly downregulated proteins (p < 0.05 and FC < 0.667); Gray : non-significant proteins (p ≥ 0.05 or 1/1.5 ≤ FC ≤ 1.5).
      3. Click Generate Image and download the high-resolution image (PDF/SVG format) for publication.
  2. Hierarchical clustering heatmap
    1. Navigate to the tool at http://www.omickits.com/open/tooldetail?id=17 
      1. Access the clustering heatmap tool.
      2. Upload the filtered DEP expression matrix.
      3. Set the following parameters: Normalization method: row-wise Z-score to eliminate scale differences; Distance metric: Euclidean distance; Clustering method: complete linkage hierarchical clustering; Optional: enable column and/or row clustering depending on sample grouping.
      4. Click Run to generate the heatmap.
      5. Download and save the heatmap as a publication-ready image.
        NOTE: The heatmap visually represents the similarity and divergence of protein expression patterns across samples.
  3. GO functional annotation and enrichment analysis
    1. Install and load required R packages:
      library(clusterProfiler)
      library(org.Hs.eg.db)

      library(ggplot2)
    2. Convert protein IDs (e.g., Uniprot or gene symbols) into Entrez IDs:
      entrez_ids <- bitr(diff_proteins, fromType = "UNIPROT", toType = "ENTREZID", OrgDb = org.Hs.eg.db)
    3. Perform GO enrichment analysis:
      go_enrich <- enrichGO(gene = entrez_ids$ENTREZID, OrgDb = org.Hs.eg.db, keyType = "ENTREZID", ont = "BP")
    4. Visualize the results using dot plots:
      1. dotplot(go_enrich, showCategory = 20)
        Formula:
        Rich Factor = (a/b) / (c/d)
        Where:
        a = number of DEPs annotated to the term;
        b = total number of DEPs;
        c = number of background proteins annotated to the term;
        d = total number of background proteins.
  4. KEGG pathway annotation and enrichment analysis
    1. Conduct KEGG enrichment analysis:
      kegg_enrich <- enrichKEGG(gene = entrez_ids$ENTREZID, organism = "hsa")
    2. Visualize KEGG pathway results:
      barplot(kegg_enrich, showCategory = 20)
    3. Customize plots using ggplot2 for publication formatting.

6. Initial screening of proteins using ROC curve analysis

  1. Data preparation : Load the proteomics dataset containing all differentially expressed proteins (DEPs) identified from the autism spectrum disorder (ASD) and control groups. Ensure the dataset includes protein expression values for both groups, with clear labels indicating ASD and control samples.
  2. Perform ROC curve analysis.
    1. Use the pROC package in R to conduct Receiver Operating Characteristic (ROC) curve analysis for each protein.
    2. Evaluate the ability of each protein to distinguish between the ASD and control groups by calculating the Area Under the Curve (AUC).
      AUC = 0.5: No discrimination (equivalent to random chance).
      0.7 ≤ AUC < 0.8: Acceptable discrimination.
      0.8 ≤ AUC < 0.9: Excellent discrimination.
      AUC ≥ 0.9: Outstanding discrimination.
      NOTE: The AUC represents the probability that a randomly selected individual from the ASD group has a higher protein level than a randomly selected individual from the control group. A higher AUC indicates better diagnostic performance, with values above 0.8 generally considered clinically meaningful in biomarker studies.
    3. Record the AUC values for all proteins.
  3. Select candidate biomarkers.
    1. Identify proteins with an AUC greater than 0.7 as candidate biomarkers.
    2. Export the list of candidate biomarkers for further analysis.
  4. Visualize results.
    1. Use the ggplot2 package in R to create visualizations of the ROC curves for the top-performing proteins.
    2. Include the AUC values in the plot legends for clarity.

7. Secondary screening using Random Forest

  1. Prepare input data.
    1. Use the list of candidate biomarkers obtained from the ROC analysis as input for random forest analysis.
    2. Ensure the dataset is formatted appropriately, with rows representing samples and columns representing protein expression values.
  2. Train the Random Forest model.
    1. Apply the random forest algorithm using the randomForest package in R.
    2. Set the number of trees (ntree) to 500 and the number of variables randomly sampled at each split (mtry) to the square root of the total number of features.
    3. Assess feature importance using the MeanDecreaseAccuracy metric, which measures the reduction in model accuracy when a specific feature is removed.
    4. Train a random forest model using the randomForest package in R:
      R. library(randomForest)
      # Example: predicting group (e.g., ASD vs. control) using protein levels

      rf_model <- randomForest(x = protein_data,
      y = as.factor(group),
      importance = TRUE, # Required to compute feature importance
      ntree = 500) # Number of trees
    5. Extract feature importance metrics using the importance() function:
      R. importance_scores <- importance(rf_model)
    6. Retrieve the MeanDecreaseAccuracy values and sort them in descending order:
      R. mean_dec_acc <- importance_scores[ , "MeanDecreaseAccuracy"]
      importance_rank <- sort(mean_dec_acc, decreasing = TRUE)
    7. Visualize feature importance using the built-in varImpPlot() function:
      R. varImpPlot(rf_model, main = "Feature Importance (Mean Decrease in Accuracy)")
      NOTE: The MeanDecreaseAccuracy metric reflects how essential each feature is to the model's predictive performance. A large decrease in accuracy upon removal indicates high importance. This approach is particularly useful for biomarker discovery, as it helps prioritize proteins or genes with the strongest discriminatory power between groups.
    8. Export the importance scores for reporting or downstream analysis:
      R. importance_table <- data.frame(
      Feature = names(importance_rank),
      MeanDecreaseAccuracy = importance_rank
      )
      write.csv(importance_table, "feature_importance.csv", row.names = FALSE)
    9. Rank the proteins based on their MeanDecreaseAccuracy scores.
    10. Select the top 15 proteins with the highest MeanDecreaseAccuracy scores as the most significant features for subsequent modeling.
    11. Export the list of these proteins for further validation.
      NOTE: Proteins with low MeanDecreaseAccuracy values may have minimal impact on model performance if removed.
    12. Highlight the biological relevance of the selected proteins, particularly those related to immune functions or pathways implicated in ASD.

8. Combine results for final biomarker selection.

NOTE: Ensure R is installed with the following packages: pROC, randomForest, and ggplot2. Ensure that the proteomics dataset is preprocessed and normalized before analysis. Save the lists of candidate biomarkers and visualization plots as separate files for reference.

  1. Integrate findings.
    1. Cross-reference the results from the ROC analysis and random forest screening to identify overlapping proteins.
    2. Prioritize proteins that appear in both analyses as highly reliable candidate biomarkers.
    3. Perform additional validation steps, such as leave-one-out cross-validation (LOOCV), to confirm the robustness of the selected biomarkers.
    4. Use logistic regression models to evaluate the predictive accuracy of the combined biomarker set.
    5. Create ROC curves and precision-recall plots for the final set of biomarkers using the ggplot2 package.
    6. Include metrics such as AUC and precision-recall values to demonstrate the diagnostic potential of the selected biomarkers.

9. Bidirectional feature selection

  1. Prepare data and define the model.
    1. Load the dataset containing protein expression values and corresponding labels (e.g., ASD vs. control). Ensure the dataset is preprocessed and normalized.
    2. Define the initial model: Use a generalized linear model (GLM) with a binomial family for classification.
    3. Use AIC as the evaluation metric to compare models during feature selection.
  2. Perform forward feature selection.
    1. Start with an empty model containing only the intercept term.
    2. Add one feature at a time based on the largest reduction in AIC.
    3. Record the AIC value after each addition. Stop when no further reduction in AIC is observed.
  3. Perform backward feature selection.
    1. Train a model using all available features.
    2. Remove one feature at a time based on the smallest increase in AIC.
    3. Record the AIC value after each removal. Stop when no further reduction in AIC is observed.
    4. Combine forward and backward steps.
  4. Alternate between forward and backward selection.
    1. Perform one round of forward feature selection, followed immediately by one round of backward feature selection. Repeat this process until no further improvements in AIC are observed.
    2. Alternative approach: Start with backward feature selection, then perform forward feature selection. Evaluate the effect of adding previously removed features back into the model.
  5. Finalize the selected features.
    1. Export the final list of selected features and their corresponding coefficients (Supplemental Figure S1).

10. Cross-validation of bidirectional feature selection using logistic regression with leave-one-out method

NOTE: Ensure R is installed with the following packages: caret, pROC, and ggplot2. The proteomics dataset should be preprocessed and normalized before analysis. Save the confusion matrix, ROC curve, and model summary as separate files for reference.

  1. Prepare the data and define the model.
    1. Load the dataset containing protein expression values and corresponding labels (e.g., ASD vs. control) from the GLMSTEP/bothFitModel.txt file. Ensure the dataset is preprocessed and normalized.
    2. Define the initial model using a generalized linear model (GLM) with a binomial family for classification.
    3. Use accuracy and Kappa coefficient as evaluation metrics to assess model performance during cross-validation.
  2. Perform leave-one-out cross-validation.
    1. Initialize cross-validation using the caret package in R to implement leave-one-out cross-validation (LOOCV).
    2. Fit the logistic regression model using the eight selected features.
    3. Record the accuracy and Kappa coefficient for each iteration of cross-validation.
  3. Analyze the cross-validation results.
    1. Summarize results.
      NOTE: The results of the LOOCV process will look like (as in this study): Generalized Linear Model, 169 samples, 8 predictors, 2 classes: 'A', 'B', Resampling: Leave-One-Out Cross-Validation , Summary of sample sizes: 168, 168, 168, 168, 168, 168, ... , Resampling results: Accuracy Kappa 0.9526627 0.9024531.
    2. Interpret the metrics.
      NOTE: Here, the model achieved an accuracy of 0.9527 and a Kappa coefficient of 0.9025 , indicating excellent agreement between predicted and observed outcomes.
      1. Look at the Kappa coefficient to gauge the model's predictive power. The Kappa coefficient ranges from -1 to 1, where 0 indicates random prediction and 1 indicates perfect agreement.
        NOTE: In this study, the Kappa value of 0.9025 reflects the model's strong predictive power.
  4. Evaluate the model coefficients.
    1. Examine the coefficients of the logistic regression model to understand the contribution of each feature. Evaluate the null deviance, residual deviance, and AIC to confirm the model's fit.
      NOTE: For example, in this study, we got Null deviance: 2.2928e+02 on 168 degrees of freedom, Residual deviance: 2.2378e-07 on 160 degrees of freedom, AIC: 18, Number of Fisher Scoring iterations: 25.
  5. Visualize the results.
    1. Create a confusion matrix to visualize the predictive performance of the model.
    2. Plot the Receiver Operating Characteristic (ROC) curve to evaluate the model's classification performance.
    3. Interpret the result s. Calculate the area under the curve (AUC) to obtained the classification performance index of the model.
      ​NOTE: The ROC curve demonstrates the trade-off between true positive rate and false positive rate. The area under the curve (AUC) should be close to 1, indicating excellent classification performance. The ROC curve reflects the changes in the true positive rate and false positive rate of the model at different thresholds. The larger the AUC value, the better the performance of the model.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The study included 99 children with ASD and 70 age-matched controls (3-7 years), with balanced sex distribution (Supplemental Table S2). Serum was collected after overnight fasting using standardized protocols: blood was drawn into serum separator tubes, allowed to clot at room temperature for 30 min, then centrifuged at 1,500 × g for 10 min at 4 °C. The supernatant was aliquoted and stored at −80 °C until further processing. High-abundance proteins (e.g., albumi...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The protocol described in this manuscript outlines a comprehensive approach for identifying immune-related molecular biomarkers in autism spectrum disorder (ASD) using data-independent acquisition (DIA) mass spectrometry and machine learning techniques. The important steps within the protocol ensure reliable and reproducible results, while also highlighting areas where modifications or troubleshooting may be necessary (Table 2).

One of the key steps in the protocol is the coll...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no conflicts of interest to declare.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Thank you to all the members of the central laboratory and those who have helped with this project.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Reagents and ChemicalsAcetonitrile (HPLC Grade)Fisher ScientificA18-50
Reagents and ChemicalsAmmonium bicarbonate (NH?HCO?)Sigma-Aldrich38939
Reagents and ChemicalsAmmonium formateSigma-Aldrich90265
Reagents and ChemicalsBovine Serum Albumin (BSA)Thermo Fisher Scientific23212
Reagents and ChemicalsDithiothreitol (DTT)Sigma-Aldrich43815
Reagents and ChemicalsFormic acid (0.1%)Thermo Fisher Scientific28905
Reagents and ChemicalsIodoacetamide (IAA)Sigma-AldrichI1149
Reagents and ChemicalsMethanol (HPLC Grade)Fisher ScientificA452-4
Reagents and ChemicalsTrifluoroacetic acid (TFA)Sigma-AldrichT6508
Reagents and ChemicalsUreaSigma-AldrichU5378
Kits and Specialized ReagentsBCA Protein Assay KitThermo Fisher Scientific23227
Kits and Specialized ReagentsC18 Sep-Pak CartridgesWatersWAT023590
Kits and Specialized ReagentsC18 StageTips (homemade)3M Empore™
Kits and Specialized ReagentsHigh-Abundance Protein Depletion KitMillipore Sigma122642
Kits and Specialized ReagentsiRT Standard PeptidesBiognosys AG
Kits and Specialized ReagentsLysozyme ELISA KitWuhan Fine Biotech Co. Ltd.
Kits and Specialized ReagentsTrypsin/LysC Enzyme MixPromegaV5071
EquipmentCentrifugeEppendorf5430R
EquipmentEasy-nLC 1200 SystemThermo Fisher Scientific
EquipmentNanodrop One SpectrophotometerThermo Fisher ScientificND-ONE-W
EquipmentQ Exactive HF-X Mass SpectrometerThermo Fisher Scientific
EquipmentSpeedVac ConcentratorThermo Fisher ScientificSPD131DDA
EquipmentSwinning-Bucket Rotor CentrifugeVarious
EquipmentWaters XBridge BEH130 ColumnWatersC18, 3.5 μm, 2.1×150 mm
EquipmentAgilent 1260 HPLC SystemAgilent1260 Infinity II
Software and Online ToolsBioconductor (R packages)bioconductor.org
Software and Online Toolscaret (R package)CRANcaret_6.0-93
Software and Online ToolsclusterProfiler (R package)Bioconductor4.0.5
Software and Online ToolsDIA-NNDIA-NN softwarev1.8
Software and Online Toolsggplot2 (R package)CRAN3.4.0
Software and Online ToolsMaxQuantMax Planck Institute1.6.17
Software and Online Toolsomickits.comOmiKits Cloud Platformhttp://www.omickits.com
Software and Online ToolspROC (R package)CRAN1.18.0
Software and Online ToolsrandomForest (R package)CRAN4.7-1.1
Software and Online ToolsSpectronaut Pulsar XBiognosys AG17
Software and Online ToolsUniProtKB Human Databaseuniprot.orgRelease 2019_10
Other MaterialsLow-bind Microcentrifuge TubesEppendorf30120094
Other MaterialsSerum Separator Tubes (SST)BD Biosciences367988
Other Materials3M Empore™ C18 Disks3M

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Ophir, Y., Rosenberg, H., Tikochinski, R., Dalyot, S., Lipshits-Braziler, Y. Screen time and autism spectrum disorder: A systematic review and meta-analysis. JAMA Netw Open. 6 (12), e2346775(2023).
  2. He, S., Zhou, F., Tian, G., Cui, Y., Yan, Y. Effect of anesthesia during pregnancy, delivery, and childhood on autism spectrum disorder: A systematic review and meta-analysis. J Autism Dev Disord. 54 (12), 4540-4554 (2024).
  3. Alkhonezan, S. M., Alkhonezan, M. M., Alshayea, Y., Bukhari, H., Almhizai, R. Factors influencing the lives of parents of children with autism spectrum disorder in saudi arabia: A comprehensive review. Cureus. 15 (11), e48325(2023).
  4. Larson, C., Thomas, H. R., Crutcher, J., Stevens, M. C., Eigsti, I. M. Language networks in autism spectrum disorder: A systematic review of connectivity-based fmri studies. Rev J Autism Dev Disord. 12 (1), 110-137 (2025).
  5. Kodak, T., Bergmann, S. Autism spectrum disorder: Characteristics, associated behaviors, and early intervention. Pediatr Clin North Am. 67 (3), 525-535 (2020).
  6. Zwaigenbaum, L., et al. Early intervention for children with autism spectrum disorder under 3 years of age: Recommendations for practice and research. Pediatrics. 136 (Suppl 1), S60-S81 (2015).
  7. Whitehouse, A. J. O., et al. Effect of preemptive intervention on developmental outcomes among infants showing early signs of autism: A randomized clinical trial of outcomes to diagnosis. JAMA Pediatr. 175 (11), e213298(2021).
  8. Ngounou Wetie, A. G., et al. A pilot proteomic analysis of salivary biomarkers in autism spectrum disorder. Autism Res. 8 (3), 338-350 (2015).
  9. Mota, F. S. B., et al. Potential protein markers in children with autistic spectrum disorder (asd) revealed by salivary proteomics. Int J Biol Macromol. 199, 243-251 (2022).
  10. Corbett, B. A., et al. A proteomic study of serum from children with autism showing differential expression of apolipoproteins and complement proteins. Mol Psychiatry. 12 (3), 292-306 (2007).
  11. Hu, E., et al. Data independent acquisition proteomics and machine learning reveals that proteins associated with immunity are potential molecular markers for early diagnosis of autism. Clin Chim Acta. 573, 120238(2025).
  12. Shen, L., et al. Biomarkers in autism spectrum disorders: Current progress. Clin Chim Acta. 502, 41-54 (2020).
  13. Yap, C. X., et al. Autism-related dietary preferences mediate autism-gut microbiome associations. Cell. 184 (24), 5916-5931.e17 (2021).
  14. Zhang, H., et al. The use of data independent acquisition based proteomic analysis and machine learning to reveal potential biomarkers for autism spectrum disorder. J Proteomics. 278, 104872(2023).
  15. Francis, K., et al. Brain-derived neurotrophic factor (bdnf) in children with asd and their parents: A 3-year follow-up. Acta Psychiatr Scand. 137 (5), 433-441 (2018).
  16. Bjorklund, G., et al. The impact of glutathione metabolism in autism spectrum disorder. Pharmacol Res. 166, 105437(2021).
  17. Bao, X. H., et al. Olink proteomics profiling platform reveals non-invasive inflammatory related protein biomarkers in autism spectrum disorder. Front Mol Neurosci. 16, 1185021(2023).
  18. Ngounou Wetie, A. G., et al. Comparative two-dimensional polyacrylamide gel electrophoresis of the salivary proteome of children with autism spectrum disorder. J Cell Mol Med. 19 (11), 2664-2678 (2015).
  19. Suganya, V., Geetha, A., Sujatha, S. Urine proteome analysis to evaluate protein biomarkers in children with autism. Clin Chim Acta. 450, 210-219 (2015).
  20. Shen, L., et al. Itraq-based proteomic analysis reveals protein profile in plasma from children with autism. Proteomics Clin Appl. 12 (3), e1700085(2018).
  21. Shen, L., et al. Proteomics study of peripheral blood mononuclear cells (pbmcs) in autistic children. Front Cell Neurosci. 13, 105(2019).
  22. Mesleh, A., et al. Blood proteomics analysis reveals potential biomarkers and convergent dysregulated pathways in autism spectrum disorder: A pilot study. Int J Mol Sci. 24 (8), 7443(2023).
  23. Yang, J., et al. Association between plasma proteome and childhood neurodevelopmental disorders: A two-sample mendelian randomization analysis. EBioMedicine. 78, 103948(2022).
  24. Wisniewski, J. R., Zougman, A., Nagaraj, N., Mann, M. Universal sample preparation method for proteome analysis. Nat Methods. 6 (5), 359-362 (2009).
  25. Wu, Y., et al. Quantitative proteomics analysis of serum and urine with dia mass spectrometry reveals biomarkers for pediatric obstructive sleep apnea. Arch Bronconeumol. 61 (2), 67-75 (2025).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Autism Spectrum DisorderImmune BiomarkersData Independent AcquisitionProteomicsMachine LearningSerum ProteomeProtein BiomarkersDifferential Protein ExpressionFunctional EnrichmentLogistic Regression

Related Articles