@‌‎n‌‎o​‌‎r‌s​‎i​‎v​‌‎a​‌eb‌‎ T​w​‎eets​‌‎

| 2022 | 2021 | 2020 | 2019 | 2018 |
Mon Oct 31 16:16:46 +0000 2022PXD033168: there are the many that talk about clinical applications & then there are the few actively pursuing them.
Mon Oct 31 14:50:44 +0000 2022PXD032110: most of the samples have good signals from pretty much all of the proteins characteristic of azurophilic granules, like BPI, ELANE & MPO.
Mon Oct 24 18:01:47 +0000 2022PXD029424: another good example of how analyzing gel bands is still a viable sub-genre in the field.
Mon Oct 17 15:49:41 +0000 2022@Smith_Chem_Wisc PXD005629 or PXD002516 for all criteria PXD010138 doesn't quite fit the criteria, but may be worth a look. PXD014606 if you want to compare with low res fragment ion data set.
Sun Oct 16 22:41:29 +0000 2022PXD035020: mouse milk proteomics is pretty rare.
Thu Oct 06 14:31:47 +0000 2022PXD032003: another good example of the pervasive infestation of XMRV in lab mice. At this point, it is close to being an endogenous retrovirus in the little critters.
Wed Oct 05 16:36:43 +0000 2022PXD034897 has an unusually prominent pattern of peptide isoelectric points in each of the LC/MS runs, because of the way the high pH fractions were mixed.
Tue Oct 04 15:08:30 +0000 2022PXD029659 has some very nice examples of protein data from the always unexpected Parainfluenza virus 5.
Mon Sep 26 19:11:50 +0000 2022PXD019869: Welcome to Collagen City! 🐭
Mon Sep 26 17:48:27 +0000 2022PXD031451: sounds kind of interesting. I don't remember anyone using TMT labeling with MHC class I peptides before.
Mon Sep 26 17:34:41 +0000 2022PXD034970: if you really want to see how well you can do using not-the-exactly-right-strain (or species) genomes to interpret data, this one looking at S. aureus antibiotic resistant strains might be a good one to use for testing ideas.
Sun Sep 25 12:43:21 +0000 2022PXD024417: better than expected examples of class I and II MHC peptide data. Really shows off how the class II peptides reflect the extracellular environment (in this case FBS) and an eccentric collection of intracellular proteins.
Fri Sep 23 18:21:39 +0000 2022Is there some technical reason that pulldown studies like PXD030720 do not reduce & block cysteines as part of the sample workup? 🔗
Tue Sep 20 21:03:19 +0000 2022PXD031455: an étude of just how deep into the carbamylated amine background an Eclipse can take you.
Tue Sep 20 15:38:59 +0000 2022PXD028509: an interesting slice of outer membrane proteins from the always entertaining H. pylori.
Fri Sep 16 01:51:13 +0000 2022PXD036085: if you've ever wondered how many different keratins you can detect in a single sample, this one is for you. Also has some nice examples of endogenous R+deimination (citrulline formation).
Mon Sep 12 23:42:41 +0000 2022PXD032089: some nice spermatocyte phosphorylation data. It is a rarely sampled cell type.
Fri Sep 09 13:46:00 +0000 2022PXD036334: for anyone interested in using chromatographic separation effects to improve PSM id algorithms, this one has some very stark examples.
Tue Sep 06 00:17:15 +0000 2022PXD026371: nice data & just like PXD032355, pancreatic ductal adenocarcinoma tissue has lots of collagen
Mon Sep 05 01:51:33 +0000 2022PXD032042: provides an interesting look at L-A & L-BC virus levels in multiple strains of S. cerevisiae.
Sun Sep 04 16:09:54 +0000 2022PXD032355: nice data that really shows how big the collagen signals can be in cancer tissue samples (in this case ovarian clear cell carcinoma).
Sat Sep 03 18:11:52 +0000 2022PXD025401 has some pretty top-drawer Trypanosoma brucei data. If you are into kinetoplastids (& who isn't), it may be just the amuse-bouche you needed to get interested in proteomics.
Sat Sep 03 17:25:44 +0000 2022PXD018214 is kind of interesting. It is unusually enriched in PSMs with multiple S/T+phosphoryl on the same peptide. Worth a look if that sort of thing floats your boat, but old-timey biochemists might want to give it a miss.
Thu Aug 18 12:44:30 +0000 2022PXD032877: also has some nice phospho-peptide enrichment from tissue (liver).
Thu Aug 18 01:20:34 +0000 2022PXD027693: some nice phospho-peptide enrichment from lung tissue samples.
Tue Aug 16 00:18:29 +0000 2022PXD034933: Keratin (proteomics' great Shaitan) you win again. Well played.
Mon Aug 15 20:06:28 +0000 2022PXD022605: interesting.
Tue Aug 09 18:50:03 +0000 2022PXD024491 is kind of a rarity: glam journal but interesting data.
Wed Jul 27 15:30:59 +0000 2022PXD035524 has good examples of the products of the side reaction that generates N-succinylation, which can occur as an unintended result of TMT peptide derivatization (~14% of assignable PSMs).
Sat Jul 23 13:23:01 +0000 2022Like CSN1S1:p, the only real exception to the normal tissue distribution of CSN2:p is PXD020222. But that particular lab frequently generates exceptional tissue distributions, so it should be taken with a grain of salt (& maybe a shot of schnapps).
Fri Jul 22 14:23:01 +0000 2022@ProtifiLlc Human CSN1S1:p is only really found in milk, with the odd exception of PXD020222, where it is found in many affinity purification expts from A-549 cells. Don't know why.
Thu Jul 21 18:21:18 +0000 2022PXD035324 has some excellent examples of how much fetal calf serum affects MHC class 2 peptide experiments, while it is undetectable in MHC type 1.
Sun Jul 17 15:22:41 +0000 2022PXD031078 has some pretty nice timsTOF Staphylococcus aureus data.
Sat Jul 16 19:40:15 +0000 2022If you want to look for interesting features in mouse blood plasma, PXD030303 is a good place to start.
Fri Jul 15 01:12:57 +0000 2022PXD033413 has some nice high resolution data from mitochondria preparations obtained using human retinal pigment epithelium samples.
Tue Jul 05 18:06:53 +0000 2022PXD034264 exemplifies the pros & cons of using cell lines & metal oxide affinity purification methods with LC-MS/MS to identify phosphorylation acceptor sites.
Sat Jul 02 17:31:29 +0000 2022PXD029730 has really grown on me. Nice work.🧐
Wed Jun 29 15:49:55 +0000 2022If you want a high-res data set from a human cell line (MHCC97H) to look for SAAVs, PTMs detectable without purification (K+GGyl, S/T/Y+phospho, R+dimethyl), viruses or which FBS proteins you might expect to find, PXD032201 would be a good choice.
Wed Jun 29 15:49:29 +0000 2022Normally I find data sets constructed specifically for bioinformatics purposes to be not so great, but PXD032201 is the exception that proves the rule.
Wed Jun 29 00:48:01 +0000 2022PXD029368 has some pretty good examples of iodo-tyrosine PSMs from R. rattus tissue samples.
Sun Jun 26 19:18:59 +0000 2022PXD029747 has some good examples of the human herpesvirus 4 (Epstein-Barr virus) proteins you can expect to find in samples derived from lymphoblastoid cell lines (LCLs)
Sat Jun 25 12:12:04 +0000 2022@astacus @LewisGeer Not a great way to put it. I use deamidation as an indicator of how well a set of PSM assignment parameters is working by monitoring the ratio N:Q deamidation. For PXD022435, that ratio is about 6:1 for the phosphopeptide data & 4:1 for the proteome data.
Sat Jun 25 00:14:53 +0000 2022PXD034877: blast from the past or just some closet cleaning?
Fri Jun 24 18:25:46 +0000 2022@LewisGeer If you really want a challenge, I'd recommend the phosphopeptide data from PXD022435. It has it all: superSILAC, low level carbamylation, easy to distinguish deamidation, modest amounts of M & W oxidation and good strong signals from iRT chromatography calibrant peptides.
Fri Jun 24 01:53:46 +0000 2022Amusing (at least to me). In PXD033584, the PubMed ID is 1 digit short, pointing to 🔗
Tue Jun 14 19:29:22 +0000 2022First Monkeypox data set I've seen since 2007―PXD034494 from the Jean Armengaud mob.
Fri May 13 17:38:37 +0000 2022If you want to test a general purpose PTM-discovery PSM assignment algorithm & need a big volume of very challenging data, PXD022435 is pretty much tailor-made for the purpose.
Wed May 11 15:19:30 +0000 2022If you are interested in the pathogen Pseudomonas aeruginosa (commonly observed in urine samples), PXD025827 has some excellent examples of what you should be seeing in clinical isolates.
Tue May 10 15:26:45 +0000 2022If you are interested in the human pathogen Proteus mirabilis & want to know what you can see using proteomics methods, examining PXD031222 would be a good place to start.
Wed Apr 27 12:10:40 +0000 2022If you need some protein phosphorylation data to work with & you have grown bored of mammalian samples, PXD032260 is the best C. elegans TiO₂ experiment I've seen in quite a while.
Tue Apr 26 00:37:05 +0000 2022PXD010700: once again the Armengaud group does not disappoint.
Fri Apr 15 15:11:06 +0000 2022Credit: PXD028931, made available supporting 🔗
Sun Mar 27 02:03:45 +0000 2022I'm only a few experiments in, but PXD032710 is looking like pretty high quality data from an interesting set of tissue samples.
Mon Mar 21 19:34:18 +0000 2022If you are interested in finding some hydroxyproline-containing peptides in cell culture data where the collagen signals aren't too overwhelming, PXD031053 looks like good data to try out. It has some issues (IAA & urea artifacts), but nothing out of the ordinary.
Sun Mar 06 19:20:59 +0000 2022@neely615 @nesvilab It appears to be an error in the link. The FTP link in PRIDE is to /pride/data/archive/2022/01/PXD023175 but the data is at /pride/data/archive/2022/02/PXD023175
Mon Feb 28 20:23:23 +0000 2022PXD020343 has good examples of what you could expect to detect in human platelet preparations using single, long HPLC runs.
Wed Feb 23 15:42:19 +0000 2022SCP finally moves into the "Ultra" era: PXD024043.
Tue Feb 22 03:31:22 +0000 2022PXD024807 isn't so bad, either. Different, but not bad.
Tue Feb 22 01:08:43 +0000 2022As phosphorylation data goes, PXD031405 is pretty good.
Tue Feb 15 17:51:37 +0000 2022PXD030591 is a great example data set demonstrating how highly represented blood plasma proteins can be in HLA-DP presented class 2 peptides, even when the those proteins are from bovine FBS.
Fri Feb 11 16:44:39 +0000 2022@statesdj @Dereklowe None, as far as I can remember, but I will check. The only one I've seen in ProteomeXchange is PXD020112 (Phoca largha), but when I looked it didn't seem to have an annotated proteome available.
Thu Feb 10 17:23:21 +0000 2022If you want a TMTpro dataset to fool around with, PXD024513 (🔗) is worth a look. Nearly complete derivatization, very few detectable side reactions & >95% recovery of Cys-containing peptide. The only downside is that 10-12% of the PSMs are semi-tryptic.
Thu Feb 03 19:01:04 +0000 2022PXD020972 has some nice high resolution MS/MS runs (from 🔗).
Mon Jan 31 14:55:15 +0000 2022PXD024716 (from 🔗) has some great examples of side-reaction peptide N-terminal succinylation in TMT10 preps, particularly how that modification segregates in high pH RP chromatography.
Fri Jan 21 01:36:02 +0000 2022PXD029418 has some examples of good recovery of cysteine-containing-peptides with the cysteines converted to the corresponding sulfonic acid derivative.
Wed Jan 19 16:56:24 +0000 2022PXD010504: if you are interested in pathological bacteria, this is a nice set of data from 🔗
Sun Jan 16 15:12:09 +0000 2022PXD022884 from 🔗 has some spiffy high resolution HLA class I peptide data
Sat Jan 15 18:43:26 +0000 2022PXD025547 (M. Bruschi, et al., 🔗) has some great examples of bacterial proteins from patients with undiagnosed urinary tract infections.
Thu Jan 13 16:30:59 +0000 2022The FLAG purified samples in PXD026306 have the interesting feature that you can see the phosphorylation state of the bait (MAPT) in the data
Wed Jan 12 14:21:02 +0000 2022If you want a high quality dataset with relatively large, consistent single shot HPLC runs (95k per) from human samples with very few experimental artifacts, PXD025174 is worth a look.
Wed Dec 22 17:05:04 +0000 2021@lukas_k Average is difficult to find in proteomics: almost every data set tries out some novel protocol. But, if you want: MHC 1/2 peptides: PXD028633 Phosphopeptides: PXD027198 Lys-C (+8 SILAC): MSV000086426 Vanilla trypsin: PXD027258 Trypsin +TMT10: PXD023979
Wed Dec 22 16:55:52 +0000 2021@lukas_k If you don't mind low res fragment masses, PXD016924 has quite a bit going for it wrt testing. Good, consistent chromatography & MS/MS applied to multiple proteolytic enzymes.
Thu Dec 16 16:01:49 +0000 2021PXD021331: while I always like oocyte data sets, the associated manuscript has a title any proteome-geek will love. 🔗
Tue Dec 14 16:26:26 +0000 2021If you are interested in mouse urine, PXD030347 is a good data set to look at. It has quality runs and several of the individuals show strong signals from UTIs: the strongest, HFX6_P6.raw, identifies >40 Proteus mirabilis proteins (out of 405 total).
Sat Dec 11 13:50:09 +0000 2021The data, PXD028633, shows that these methods can be used to prepare high quality class 1 & 2 samples that have nearly theoretical physical properties. If you want to either make samples or have some excellent exemplars for data analysis, this study is a good place to start.
Thu Dec 09 13:52:37 +0000 2021PXD027198 is an interesting data set if you are interested in phosphorylation acceptor site analysis: made available as part of 🔗
Wed Dec 01 20:06:58 +0000 2021PXD027258: while I normally won't get out of bed for something that is anything less than "ultra", this mere "super" is kind of interesting.
Wed Dec 01 15:33:14 +0000 2021e.g., the % of PSMs by C-terminal AA for ELITE-RSLC012686.raw in PXD016924: AA % of PSMs A 0.5 C 0.5 D 0.5 E 0.07 F 25.7 G 0.2 H 2.0 I 0.2 K 1.6 L 34.0 M 2.7 N 1.1 P 0.2 Q 1.2 R 1.3 S 0.5 T 0.6 U 0 V 0.3 W 6.2 Y 20.8 🔗
Mon Nov 29 15:41:00 +0000 2021PXD016924 combines some credible mitochondrial isolations with different protease treatments. Trypsin, Asp-N, Glu-C & chymotrypsin digests are done, in batches of 128 runs for each protease. May be particularly useful for the AI-prediction-of-RT crowd.
Thu Nov 18 14:50:53 +0000 2021If you are interested in the problems & benefits of the systematic study of proteins obtained from clinical tissue samples, PXD024124 is a really good resource 🔗 (⭐️⭐️⭐️⭐️)
Wed Nov 17 15:47:19 +0000 2021The files with names containing "Ross_Birsoy_SPS-MS3_TMT" in PXD027673 give an unusually detailed peak inside of human mitochondria & use the TMTpro quantitation reagent.
Sun Nov 14 18:04:39 +0000 2021@ypriverol @slavov_n @BiswapriyaMisra @MetabolomicsWB @MetaboLights I don't know about the DIA dataset (PXD002952), but the other 2 are really eccentric choices for popularity.
Tue Nov 09 17:37:04 +0000 2021I don't know what about their protocol makes it so, but PXD023979 did a great job of TMT10 derivatization.
Wed Nov 03 14:20:54 +0000 2021PXD004652 has some great data if you are interested in using isoelectric focusing as a method of peptide separation. Does a good job of producing reproducible analyses from multiple samples.
Sun Oct 31 15:54:49 +0000 2021PXD025621 does a great job of K+succinyl enrichment from murine brain tissue.
Sat Oct 30 13:22:50 +0000 2021If you want to use a TMT-based dataset, but need one where the derivatization was done really well (very few artifacts), PXD021038 is a pretty good exemplar.
Sat Oct 23 13:04:42 +0000 2021@leprevostfv It is all public, but there is quite a bit of it (~860,000 PSMs from 64,632 LC-MS/MS runs). Did you want a list of the PXD/Massive/TRANCHE accessions?
Wed Oct 06 14:39:52 +0000 2021@neely615 If you are still interested, here are the obs. of Streptococcus salivarius enolase in human saliva samples 🔗 Click the metadata tab & you can scrape out the PXD numbers you find interesting. Enolase is A1 on the jukebox when Streptococcus sp. are present.
Sun Oct 03 14:07:30 +0000 2021If you an algorithm you think would be good for determining sequence specificity in proteases, the data in folder IPX0002310001 (part of PXD028891) would be pretty much ideal for testing it.
Sat Sep 25 17:16:16 +0000 2021The DDA data in PXD022198 is interesting. I don't know how, but they were unusually effective at minimizing oxidation during their sample workup, while at the same time promoting N-deamidation.
Fri Sep 24 14:13:22 +0000 2021I'm liking PXD016790 as an example of the effects you can expect to find in samples processed using the popular 8 fraction high pH MDC kit.
Sat Sep 18 12:40:50 +0000 2021@theoneamit @VATVSLPR @emwhitti Gels have retained some popularity as part of a protein-level separation MDC method, by cutting an entire gel into segments (5-10) and running each piece separately, e.g. PXD023266.
Wed Sep 15 15:18:15 +0000 2021PXD020226 is very useful data if you are interested in S/T phosphorylation in prokaryotes 🔗
Tue Sep 07 20:28:50 +0000 2021The more I look at the analysis of PXD023938, the more interesting it looks as a potential model data set.
Tue Aug 31 12:38:51 +0000 2021If anyone is interested in biotyping pathogens in UTIs, PXD005219 has some nice examples.
Tue Aug 24 00:53:45 +0000 2021It you want a few HLA class II runs to use as examples, you could do a lot worse than PXD023032.
Thu Aug 19 14:37:38 +0000 2021@FourMohrYears They are RAW files: 1 = B20171030-10_Huh7_PV3h_1.raw 2 = B20171030-10_Huh7_PV3h_2.raw from PXD024546.
Tue Aug 10 20:47:13 +0000 2021@Smith_Chem_Wisc You should be able to see it in quite a few of the files in PXD025048, for example: DT_01192021_20U_W_1.RAW. There should also be a 1 residue longer form detectable, 57-88, as well as the intact insulin A chain (90-110).
Thu Aug 05 14:58:30 +0000 2021If you want an interesting bacterial data set to look at, try PXD025119. The organism (T. halophilus) is rarely examined by proteomics, but it is of significant value in the food biz. It has a well annotated proteome & the data is good quality. From: 🔗
Sat Jul 24 11:54:51 +0000 2021@theoneamit PXD023446 is correct, but the data is stored in Massive at ftp://massive.ucsd.edu/MSV000086665
Fri Jul 23 16:20:51 +0000 2021PXD023446 is a good data set to use as a torture-test any protein inference algorithm.
Thu Jul 22 21:15:28 +0000 2021@UCDProteomics, what type of chromatography did you use for the first dimension in PXD018133?
Thu Jul 22 15:42:05 +0000 2021If you are looking for a good quality prokaryote data set using an interesting organism with a well annotated proteome that is NOT involved in disease, PXD019583 (Nitrospira moscoviensis, 🔗) would be a good choice.
Sat Jul 17 12:46:33 +0000 2021PXD024126 is officially my current recommended data set for understanding experimental effects associated with using TMTPro, using a Th*rmo instrument. 🔗
Thu Jul 15 18:49:25 +0000 2021PXD024126 is nice data for understanding experimental artifacts associated with TMTPro derivatization. High resolution MS/MS, good LC & enough of each common artifact to produce confirmatory distributions.
Wed Jul 14 15:49:16 +0000 2021Madita Brauer, et al. PXD022830 is an excellent example of adventitious cysteine blocking by 2-mercaptoethanol
Wed Jul 14 14:39:31 +0000 2021Good work Caleb Walker, et al. PXD020854. Interesting data: a few artifacts, but Y. lipolytica has so many sequence differences from any model species that it can serve as a good test for algorithms trained on those models.
Tue Jul 13 14:34:21 +0000 2021Congrats to Yujiao Yang, et al. (PXD027205). Very nicely done. ⭐️⭐️⭐️⭐️
Tue Jul 06 00:36:25 +0000 2021If you are looking for a human data set to use for teaching, with reproducible reps, multiple cell lines & lots of interesting features in the results, PXD022002 could be what you are after.
Wed Jun 09 01:45:13 +0000 2021While it wasn't the point of the study (PXD024967), the data files starting with "Urine_SARS_TP_TMT" represent the best example of urine from a Klebsiella pneumoniae UTI I've seen.
Thu May 27 18:44:09 +0000 2021Since I tend not to believe myself, I checked a bunch of new GG-remnant data (PXD024338) for evidence of this type of cleavage. I am relieved to say I couldn't find any significant evidence of K+GG at the C-terminus of any of the tryptic peptides.
Tue May 25 20:38:03 +0000 2021PXD022523: nicely done.
Tue May 25 16:41:25 +0000 2021@MiguelCos @astacus PXD015289 does a pretty good job on extracellular vesicles in urine & PXD020722 on unfractionated urine. The write ups for both ignore effects associated with UTIs that are easily detected in the data (E. coli in the former & P. aeruginosa in the latter).
Mon May 24 17:01:36 +0000 2021If you would like a very good phosphopeptide data set that uses dimethyl isotopic labels for quantitation, PXD020805 is well worth a look. The protocol used achieved > 95% enrichment of phosphopeptides.
Tue May 18 12:22:11 +0000 2021F12:p, MHC1 ligand peptides for this protein have only been observed in liver tissue (PXD019643), where the protein is synthesized.
Sun May 09 14:52:25 +0000 2021PXD025813 is one of those mysterious human cell culture experiments in which the data shows a surprisingly large number of PSMs for PIV5 proteins, even though there is no mention of PIV5 infection in the paper.
Wed May 05 02:09:34 +0000 2021@ImmunoFever Yes. It was observed in the data (PXD008127) from 🔗
Mon May 03 14:04:23 +0000 2021@astacus Ha! More mouse urine (PXD018023). C57BL/6J mice this time.🐭
Sat May 01 21:29:09 +0000 2021If PXD019643 shows me anything, it would be that if you want lots of MHC class 1 or 2 peptides, skip every other tissue & go straight to for the thymus.
Thu Apr 22 13:56:43 +0000 2021PXD019643 shows that even at scale, it is possible to do really good MHC class 1 or 2 sample preps. The peptide lengths observed look nearly theoretical. 🧐 These two distributions are from individual RAW files in their "adrenal gland" set of expts. 🔗
Sun Apr 18 14:37:20 +0000 2021PXD018369 (CSF) is another state-of-the-art data set generated using timsTOF PASEF (which also wins the "worst acronym in proteomics" award).
Fri Apr 16 19:12:34 +0000 2021It will depend on how the data works out, but PXD019643 certainly seems to be a rather ambitious MHC peptide project. The first few that I have analyzed look promising, though.
Tue Apr 13 15:12:52 +0000 2021PXD000653 is worth a look as exemplar data for developing algorithms to detect deamidation deliberately caused by PNGase F (or endogenously by AGA). It uses older technology (low resolution fragment ions, LTQ Orbitrap Velos Pro), but the results are really pretty good.
Sun Apr 11 15:23:50 +0000 2021An oldie but goodie: PXD004540. While probably not very useful for one of its original purposes because of poor coupling of TMT to peptides, it is a very good for the other: enriching peptides containing deglycosylated N-linked glycosylation. /1
Wed Mar 31 22:51:17 +0000 2021I realized that my life is certainly at a strange place, when I got mildly excited by PXD022458, because it is the first data set I've seen taken from Mus musculus urine samples.
Tue Mar 30 14:32:19 +0000 2021If you are interested in the combination of mouse phospho-peptides and TMT 6-plex, you should definitely check out PXD023803/MSV000086764. Also good for meiosis fans! 🐭
Thu Mar 25 13:05:14 +0000 2021GC:p.D432E, chr 4:g.71752617A>C, rs7041, (PXD004242 D:E = 231:785) vaf=52%, Δm=14.0157. VAF by population group: african 9%, european 58%, east asian 30%, south asian 54%, american 54%. DD has been associated with reduced vitamin D levels in older adult smokers. #ᐯᐸᐱ
Wed Mar 24 12:51:04 +0000 2021C3:p.P314L, chr 19:g.6713251G>A, rs1047286, (PXD004242 P:L = 1435:175) vaf=14%, Δm=16.0313. VAF by population group: african 1%, european 20%, east asian 0%, south asian 7%, american 9%. PL or LL carry increased risk of age-related macular degeneration. #ᐯᐸᐱ
Sun Mar 21 13:17:09 +0000 2021APOE:p.C130R, chr 19:g.44908684T>C, rs429358, (PXD004242 obs. 484×, PXD022817 obs. 32×) vaf=14%, Δm=53.0919. This SAAV & APOE:p.R176C form a polymorphic allele that can indicate risk of developing Alzheimer's disease. #ᐯᐸᐱ
Fri Mar 12 18:17:52 +0000 2021@byu_sam @pride_ebi PXD006932 has some pretty good HeLa cell lysate QE-HF data & it is newish (2017). PXD011163 is newer (2020) , uses QE or Fusion. It is a bit more complicated (you should add in gi|130416| for the ids), but it is essentially a HeLa lysate.
Thu Mar 11 21:28:01 +0000 2021"CSF_Pool_Pool_Pool_Pool_Pool_Pool_01_Plate01.raw" may be my new favorite file name (from PXD009589).
Fri Mar 05 12:46:19 +0000 2021PXD022817 is an excellent example of a state-of-the-art human plasma proteomics data set. Congrats guys: I'm going to take this one as the new standard 🔗
Thu Mar 04 15:28:12 +0000 2021👍😀 to the submitters of PXD022817. Including the MGF files really makes reusing the data easier for most non-Thermo raw file formats.
Thu Mar 04 00:41:41 +0000 2021PXD022621 is a really good example of yeast SILAC data using hi/low pH MDC separation. Excellent for use in tuning up algorithms and identification statistics for a midsized proteome. 🔗
Mon Mar 01 15:45:23 +0000 2021If you want to study the patterns of mouse hydroxylysine and hydroxyproline modifications in tissue, understanding the data in PXD019431 would be a good place to start.
Wed Feb 24 15:39:02 +0000 2021Analysis of PXD024113 & PXD024123 has reminded me how little community effort has gone in to coping with variability in immunoglobulin sequences. Canonical sequences just don't do the job.
Tue Feb 23 02:59:48 +0000 2021@ProtifiLlc @dtabb73 Brett Phinney is the expert on that subject (e.g. PXD016169). If you see adventitious sheep keratins in a sample, it is almost always from a European group. The keratins from wool don't share tryptic peptides with human skin KRT1, KRT2, KRT9 & KRT10.
Fri Feb 19 17:11:46 +0000 2021@LewisGeer You might want to try PXD006110 (Lactobacillus delbrueckii ssp. bulgaricus). There are some IAA artifacts on amino groups, but the MS/MS is great, >50% PSM assignment & the LC is well done. Excellent for detecting N or Q deamidation, too.
Fri Feb 19 15:58:46 +0000 2021I happened across PXD004280 again: I had forgotten how good it was from a purely technical point of view. If you want to test an algorithm with excellent human cell line data where the % of identified spectra is high (>70%), you should try this one.
Wed Feb 17 13:14:16 +0000 2021Congrats of Jiang Liu, et al. for the quality of PXD022673. The data is a single high/low pH multidimensional chromatography run with TMT, but is it a good one! ✅
Tue Feb 16 17:47:33 +0000 2021But that problem aside, the MS/MS and HPLC methods used for PXD021518 were particularly well done.
Fri Feb 12 18:45:06 +0000 2021PXD022851: nice data too👍 (no associated manuscript).
Fri Feb 12 17:28:23 +0000 2021PXD020908: nice data 👍🔗
Thu Feb 11 00:15:30 +0000 2021Thanks to everyone who participated. Based on the results from PXD014917, the phosphorylated PSM ratios for unfractionated human milk are approximately SPP1:CSN1S1:CSN2:CSN3 45 : 10 : 2 : 0 SPP1:p is also the dominant source of phosphopeptides in urine proteomics data.
Mon Feb 01 16:40:51 +0000 2021Does anybody know exactly which experiments in this paper correspond to the DDA RAW files in PXD017052? After reading the paper a few times it isn't clear to me where the data fits in to what is being described. 🔗
Mon Dec 21 13:21:51 +0000 2020NRXN1:p has considerable tryptic peptide overlap with DPYSL2, DPYSL3, DPYSL5, CRIMP1, NRXN2 & NRXN3, but good data (e.g. PXD004572 or PXD006109) contain unambiguous identifications.
Thu Dec 10 19:40:24 +0000 2020@CameronTFlower If you want an alternative, try PXD019909. Lots of challenges, but they are due to cell-specific protein biochemistry rather than experimental artifacts.
Thu Dec 10 16:28:44 +0000 2020Why do people still use PXD000561 as an exemplar data set in studies? It has a lot of technical problems and unless you really want to show that you can identify and cope with those problems, it isn't a good choice for general purpose use.
Tue Dec 01 19:05:19 +0000 2020I am genuinely excited to see the results of analyzing the data in PXD019258.
Tue Dec 01 17:05:04 +0000 2020Revisiting PXD014845, it seems like an ideal data set to explore sensitivity vs. selectivity for PSM id algorithms.
Mon Nov 30 18:39:34 +0000 2020IMHO, PXD020722 probably should have used single shot expts of individual urines to create the libraries instead of pooling them all and doing multidimensional chromatography. At least you would be able to sort out the effects of the various UTIs present in the pools.
Fri Nov 27 15:59:46 +0000 2020If you want some data to test your new algorithm (or maybe an open search will find them?), try the "Keratinocyte" data from PXD019909. It has multiple observations of peptides generated by this mechanism, e.g.: 5/9
Sat Nov 21 21:23:24 +0000 2020If you are feeling up for a challenge, the "dermis" data from PXD019909 (🔗) has one of the most naturally complex set of modifications I've run into recently.
Thu Nov 19 14:48:35 +0000 2020Anyone interested in SARS CoV and SARS COV2 protein-protein interactions in a cell should really take a close look at PXD020222 (🔗). While it has not received as much attention as some other studies, technically it is by far the most sophisticated.
Wed Nov 18 15:56:59 +0000 2020Could somebody pester the authors of 🔗 to release their lock on PXD012615? The paper has been out since July but the data is still not publicly available.
Thu Nov 12 15:17:41 +0000 2020@ypriverol If you don't mind mouse tissue, you have more options. For example, PXD018140 (Lui JJ Neuropharmacology, 181:108324, 🔗) has a lot of very nice high resolution spectra. It also has a phosphopeptide enrichment of > 95% with no labelling.
Mon Nov 09 14:02:55 +0000 2020From PXD021205: just the way you want data to work out in a good LC/MS/MS run. Nearly 40% id rate for most of the run, tapering down to 0% at both ends, with good retention time prediction correlations throughout. 🔗
Fri Oct 30 17:44:38 +0000 2020I nominate PXD004452's 39 fraction trypsin digest experiment (🔗) to be the "best" HeLa cell proteomics data set yet.
Tue Oct 27 14:27:12 +0000 2020Some of the clonal organiod proteome runs in PXD016582 (🔗) are excellent. Very clean sample preps & nice sharp chromatography with good recovery of R+dimethyl, S/T/Y+phosphoryl & H+diphthamide (which is rare!).
Mon Oct 19 21:58:54 +0000 2020PXD021588 is a new study, looking at protein-protein interactions between SARS CoV, SAR CoV2 and MERS viral protein baits and HEK 293T host proteins. 1/2
Mon Oct 19 18:34:47 +0000 2020@gingraslab1 does anybody there remember whether cysteine sidechain derivatization used for PXD018196? I couldn't find a mention of it in the bioRxiv PDF. The data looks like it is unmodified cysteine, with some adventitious glutathione modification.
Thu Oct 15 19:22:15 +0000 2020@MattWFoster PXD012308 & PXD018998 are both pretty good, high resolution human class II datasets.
Thu Oct 08 19:45:44 +0000 2020PXD018140 (Liu JJ, et al., Neuropharmacology. 2020 Sep 22;181:108324, 🔗) has some excellent phosphopeptide-enriched data for 3 different mouse brain tissues.
Thu Oct 01 13:59:54 +0000 2020If you are interested in doing a 3- channel SILAC study of phosphorylation (S,T & Y) in a cell line, PXD018566 is a good data set to examine. You will get a very good idea of the LOD/LOQ that can be obtained using this approach.
Tue Sep 29 14:29:42 +0000 2020Big thanks to Mikel Azkargorta, et al (PXD021140 & PXD021139) for including MGFs along with their timsTOF Pro raw data. Saves a lot of time & bandwidth for those of us primarily interested in the PSMs.
Fri Sep 25 18:00:06 +0000 2020The HLA type 1 & proteome data in PXD020011 are very good: nice job Fabio Marino, et al. Front. Immunol, 28 August 2020 🔗 The "proteome" data is excellent for testing SAAV-finding algos, once you allow for the 2% of PSMs with urea-induced carbamylation
Tue Sep 22 17:38:24 +0000 2020PXD018402, nice job! Good high-res data from C. elegans embryos, with minimal E. coli OMPs. Placentino M, et al. 🔗
Sun Sep 20 12:57:41 +0000 2020Has anybody else looked at the data from 🔗 (PXD016999)? I'm particularly interested in what people think about the use of a calibration channel created from of a mixture of many samples. It seems to cause issues, but maybe it fixes more than it makes?
Fri Sep 18 15:58:29 +0000 2020PXD021081, Jichang Huang, et al. 🔗 — nicely done. ⭐️
Thu Sep 17 15:08:07 +0000 2020PXD017766, nicely done (Navarro JF, et al. 🔗)
Tue Sep 15 23:17:43 +0000 2020The wind is now (as of 17:40) running at 146 kph gusting to 167 kph (91 mph/104 mph) & it is near the southeast edge of the eye wall: Viosca Knoll is the pink circled (A). The pressure has held at 983-985 mb since 12:30! 🔗
Sun Sep 13 14:14:07 +0000 2020The experiments resulting in PXD020078 generated some nice data. Good work Gajanan Sathe, et al. 🔗
Sat Sep 12 14:41:52 +0000 2020Does anybody know if there is an index that associates the data files in PXD016999 (A Quantitative Proteome Map of the Human Body) with the specific tissues?
Mon Sep 07 20:45:24 +0000 2020While the title of the manuscript may be a little over-the-top, the data (PXD016477) associated with 🔗 is really pretty interesting
Tue Sep 01 14:20:33 +0000 2020PXD015430 is my new favorite high resolution MS/MS phospho-peptide dataset. Does anybody know if it has been included in a publication yet?
Wed Aug 19 15:21:59 +0000 2020@ypriverol @AJ_Brenes @pride_ebi Does "data set" in this context mean "raw file" or a collection of files with 1 PXD number?
Thu Aug 13 16:37:32 +0000 2020Viruses are weird: data from "Proteome and phosphoproteome dynamics of CVB3 infected cells" (PXD011163) 🔗
Wed Aug 12 21:39:20 +0000 2020If you are interested in phospho-proteomics and would like to take some good high-res IMAC data out for a spin, PXD014832 from Bonucci M, et al. 🔗 is a great place to start ⭐️⭐️⭐️⭐️
Sat Aug 08 16:39:16 +0000 2020If you ever need some really good ubiquitinylation site (K-GG) data, PXD018743 is a particularly elegant example (🔗).
Sat Aug 08 16:00:51 +0000 2020PXD018430 is a great data set if you are interested in examples of arginine citrullination (aka deimination).
Fri Jul 17 20:51:02 +0000 2020@MattWFoster The companion data set from the same paper (PXD020222) used A-549 as annotated, with no traces of HEK-293T present in the data.
Fri Jul 17 19:25:23 +0000 2020PXD016231 (🔗) has well done proteomics data associated with plasma, salivary glands & saliva from patients with symptoms of Sjögren's syndrome. The authors have analyzed the human proteins detected to try to find biomarkers to aid in diagnosis.
Fri Jul 17 17:40:56 +0000 2020PXD020019 (🔗) is the 1st publicly available data set to demonstrate the extent of K+ubiquitinyl of SARS COV-2 proteins in a cell line infected with the virus. It also has proteome & phosphoproteome data confirming the results of earlier studies.
Tue Jul 14 15:54:26 +0000 2020Does anybody know if there is an index describing the provenance of the raw files associated with PXD019645? Given the number of experiments involved in 🔗 & very generic files names, re-analysis of most of the data isn't possible.
Tue Jul 07 18:14:26 +0000 2020If you are interested in Mycobacterium smegmatis (a relatively easy to grow, non-pathogenic model species that can stand in for much nastier mycobacteria), PXD017602 has a particularly good sampling of the proteome. Published as part of 🔗
Sat Jul 04 15:49:45 +0000 2020Perhaps I am being difficult, but was it impossible to label the data files associated with PXD015979 as being either "M" or "F", given that the title of the study is "Proteomics pinpoints alterations in grade I meningiomas of male versus female patients"?
Mon Jun 29 17:45:21 +0000 2020If you have any interest in the phosphorylation found in Toxoplasma gondii infected human cells, PXD019729 has the best public data available.
Mon Jun 29 15:11:28 +0000 2020Does anybody know if there is a master list for 🔗 (PXD019645) that describes the association of RAW files with their corresponding experiment?
Wed Jun 24 15:28:25 +0000 2020@neely615 @MattWFoster @KentsisResearch @BiswapriyaMisra @Sci_j_my @AlexHgO The argument doesn't square with the data from plasma EVs (PXD001194). In that data, both 20 S hydrolytic subunits (🔗) and 19 S regulatory subunits (🔗) are well represented.
Mon Jun 22 15:14:07 +0000 2020Nothobranchius furzeri (PXD016587): I'm not sure I buy the argument that the little guys are a good model for aging, but they are a great example of how adaptable even generic-looking vertebrates can be when there is an environment to exploit. 🔗
Fri Jun 12 17:51:22 +0000 2020PXD013724 has some very good chromatography and MS/MS, as well as describing a set of experiments that could be fairly easily reproduced in a lab tutorial/workshop
Fri Jun 05 16:27:06 +0000 2020PXD012921 is a nice addition to what is known about phosphorylation in human herpesvirus 5 (aka cytomegalovirus, HHV5 & HCMV) proteins. HHV5 has about 40 phosphoproteins, out of 190 protein coding genes.
Thu Jun 04 16:59:43 +0000 2020Does anybody know if PXD015523 has an associated publication? The data represents a very interesting list of proteins, quite different than most (all?) of the results I've seen.
Mon Jun 01 21:36:53 +0000 2020If you ever want to see the extent to which collagens can dominate a human tumour tissue sample, PXD014980 is a great illustration of this phenomenon (S. Lee, et al., 🔗).
Mon Jun 01 14:33:29 +0000 2020PXD018594 has 2 reps of a time course, with virus originally loaded at 2 levels and sampled at days 1, 2, 3, 4 & 7 post-infection.
Fri May 29 13:32:31 +0000 2020@PastelBio @ProteomicsNews PXD018804 & PXD018594 from the Armengaud group's SARS-CoV-2 paper popped up on PRIDE this morning. It will be interesting to see what they found.
Thu May 28 13:02:00 +0000 2020@PastelBio @ProteomicsNews At least the data (PXD018804) should be released soon.
Fri May 22 16:38:42 +0000 2020The data set PXD019119 (Cardozo, Res. Sq. 2020 🔗) is full of interesting results.
Mon May 18 14:06:35 +0000 2020PXD016126 is an interesting study of the salivary proteome. The variability in the salivary microbiome is particularly nicely illustrated by the data. NOTE: the LC/MS/MS has much wider parent ion mass accuracy distributions than are common for this type of instrument.
Thu May 14 14:09:31 +0000 2020Several months into the pandemic, PXD017710 remains the only publicly available data set with good signals from SARS-CoV-2 virus proteins that were obtained from an infected cell line (CACO-2, with easily observable ACE2:p).
Sat May 09 22:45:12 +0000 2020For anyone who wants to know what a really good Trypanosoma brucei data set looks like, try PXD016370. Lots of proteins of unknown function, with very good signals and chromatography.
Thu May 07 11:43:50 +0000 2020PXD013649 is a pretty good set of data for anyone interested in understanding the detection of MHC I or II type peptides associated with different HLAs, from probably the best experimental group in this field.
Thu Apr 30 16:49:58 +0000 2020@pwilmarth I forgot to mention that PXD014414 also has 8-9% non-tryptic cleavage (no chromatographic enrichment). Other than N-terminal protein processing, most of the cleavage is due to a chymotrypsin-like activity ([FYWL]-X). Also very nice for algorithm development!
Thu Apr 30 14:30:51 +0000 2020@pwilmarth PXD014414 is for some reason (prob. sample prep) depleted in S/T/Y phosphopeptides, but has a typical S:T site detection ratio (5:1), also making it good test data for accurately detecting rare (but there) PTMs. /2
Thu Apr 30 14:26:50 +0000 2020@pwilmarth The data set (PXD014414) has some additional nice properties for algorithm developers. It shows a nearly ideal ratio of N:Q deamidation (10:1), which make it useful for anyone who is interested in trying to reduce false positives for this rather tricky chemical modification /1
Mon Apr 27 16:47:44 +0000 2020PXD017030 is an interesting demo of an enzyme TrypN (C. thermophilum) that cleaves N-terminal to R or K, similar to LysargiNase (M. acetivorans). Enzyme works great, but if you try the analysis yourself, the Cys-blocking didn't work & allow for 3 missed cleavages.
Sun Apr 19 15:03:49 +0000 2020For anyone interest in Alzheimer's Disease mouse model systems, PXD017916 provides a very interesting insight into all of the other proteins associated with insoluble Aβ deposits.
Fri Apr 17 14:57:03 +0000 2020While I'm not a fan of most qLC/MS data, PXD016166 does demonstrate some interesting effects, particularly the phosphorylation study. Rinfret Robert C, et al., J Proteome Res. 2020 Apr 8, 🔗
Wed Apr 15 14:53:48 +0000 2020PXD016678 has some very good examples of the effects of carbamylation on PSM assignment and protein quantitation. The MS/MS and chromatography are both very well done, making this a good example data set for studying this particular sample preparation artifact.
Tue Apr 14 15:17:53 +0000 2020PXD017386 is an interesting study that you could probably use as a dry-lab data example (Mizukami H, et al., 2020, 🔗). Nothing high stakes about it, but lots of things that could use explaining.
Fri Apr 10 19:09:53 +0000 2020If anyone is looking for a data set to use to get started with HLA peptide analysis (both I & II) from the view point of experiment design, data analysis or QA/QC, I'd recommend giving PXD017149 a good hard look. Lots of data, clinical samples and pretty uniform quality.
Thu Apr 09 20:43:28 +0000 2020PXD014963 is the first time I've seen significant derivatization of peptide cysteine sidechains with glutathione.
Fri Apr 03 15:19:17 +0000 2020If you want to test your peptide id skills, try re-analyzing PXD014342. Tooth enamel is difficult at the best of times, but from archeological samples it is double tough. For extra points, try seeing how well they did in comparison to PXD009781.
Fri Apr 03 14:25:51 +0000 2020PXD016828 has some nice observations of thyroglobulin iodination in normal thyroid tissue as well as primary and metastatic tumors. Iodination is a complex, multi-step PTM that requires specialized intracellular vesicles, showing that the tumors retain this elaborate mechanism.
Wed Apr 01 18:28:14 +0000 2020PXD016519 used a particularly good batch of Lys-C (unfortunately no lot number in the paper).
Wed Apr 01 14:11:20 +0000 2020Something that may be of interest to virology-proteomics-types, PXD018117 contains data where the viral protease (nsp5) is in high concentration, resulting in enhanced non-tryptic cleavage in the ID'd peptides (e.g., qx017090.raw, qx017091.raw, qx017092.raw). /1
Tue Mar 31 21:16:23 +0000 2020The data (PXD018117) indicates that most of the nsp proteins (from ORF1AB) are missing the last few residues of their C-termini. No idea why, since these are constructs.
Wed Mar 25 15:22:33 +0000 2020Hopefully someone can answer this question, regarding #PXD018117. Are all of the viral protein constructs expressed simultaneously in one batch of cells or were the constructs expressed individually in separate batches? The manuscript isn't crystal clear about this point.
Tue Mar 24 14:49:08 +0000 2020Just got the analysis of PXD018117 from PRIDE 🔗 loaded into GPMDB 🔗 Uses sequences from: ENSEMBL (female human & FBS); RefSeq (SARS CoV2 & other common human viruses); and cRAP.
Mon Mar 23 21:09:53 +0000 2020Has anyone else reanalyzed PXD018117 yet?
Thu Mar 19 17:32:16 +0000 2020Compared to most MHC type I datasets, PXD013064 does an unusually good job at recovering cystine-containing peptides (Chen, et al. 🔗)👍
Thu Mar 19 16:10:03 +0000 2020If you want to do some further re-analysis of betacoronavirus proteomics data, PXD002358, PXD004716 & PXD004719 have a lots of data for MERS CoV infections
Wed Mar 18 17:36:35 +0000 2020@chrashwood And I would be remiss to neglect the CACO-2 data set that has ACE2:p with demonstrated positive viral receptor capability, all of PXD017710.
Wed Mar 18 16:42:01 +0000 2020@chrashwood If you are volunteering, PXD002121 would be a place to start, particularly the raw files prefixed: "+WP vs -WP D9 to D22 TMT set7 B"; "+WP vs -WP D9 to D22 TMT set6 B"; "+WP vs -WP D9 to D22 TMT set7 A"; "+WP vs -WP D9 to D15 TMT set4 B"; & "+WP vs -WP D9 to D15 TMT set4 A".
Mon Mar 16 15:54:07 +0000 2020NOTE: the SARS-CoV-22 proteins being featured this week were all observed in the data set PXD017710 using the RefSeq proteome sequence. The data was made available in conjunction with the preprint 🔗
Tue Mar 10 15:43:28 +0000 2020@wenbostar As a final note, if you want to see these peptides you do not need a "special" data set of HLA type-I peptides. Grab any one of the HLA type-I data sets from ProteomeXchange — e.g., PXD015957 or PXD013831 — and you'll find lots (by which I mean 1-5% of PSMs)
Sat Mar 07 16:46:12 +0000 2020PXD017614 does a good job of isolating and characterizing mitochondrial proteins. I can't locate an associated publication.
Fri Mar 06 19:08:54 +0000 2020PXD017025, nice data with unusually high id rates for human clinical samples. Reported in 🔗
Tue Feb 25 17:35:07 +0000 2020@dtabb73 There is a link to the FTP site on the ProteomeXchange page. In this case, ftp://ftp.pride.ebi.ac.uk/pride/data/archive/2019/04/PXD012703. And the directory seems to be absent for some reason.
Sat Feb 22 21:44:02 +0000 2020PXD015361 provides an interesting wrinkle on the phosphopeptide isolation process that I'd never seen before.
Fri Feb 21 21:50:40 +0000 2020PXD015957 has some nearly ideal MHC type I peptide data sets. Good work 🔗 🔗
Thu Feb 13 15:40:58 +0000 2020#PXD014644: swing-and-a-miss from a usually reliable group.
Tue Feb 11 21:20:41 +0000 2020@UCDProteomics But I should mention that many of the runs in PXD015087 with "Hela" in the file name contain consistent ids for large T antigen [Simian virus 40, gi:297591903], which is only usually found in HEK 293T cells.
Mon Feb 03 19:14:51 +0000 2020Comparing a Threadripper vs RPIB4 for PSM assignment (I promise I will stop reporting this very esoteric stuff): 227 MGF files from PXD016632; all times include loading spectra, PSM, FDR & p-value assignment + writing TSV final output and JSON metadata files. 🔗
Fri Jan 31 16:10:18 +0000 2020Latest RPI4B run stats: 874 MGF files (PXD015943), phosphopeptide enriched 19,286,342 spectra 6,499,449 PSMs (>90% phosphopeptides) 66 sec/MGF 58,406 sec total time 11.2 Joules per 1000 spectra incl. loading spectra, PSM, FDR & p-value assignment + writing TSV final output.
Thu Jan 30 16:11:49 +0000 2020So far, I am really liking #PXD015943 (🔗). It may be the best quality data set ever associated with a Nat. Biotech. publication.
Fri Jan 24 21:00:12 +0000 2020@ProteomicsNews @lstops @chrashwood I checked out the rat phosphorylation portion of #PXD014750 & the reagent clearly has worked. There is nearly complete coupling of the TMTpro to free peptide N-terminii and lysines & only a small amount of the succinylation side reaction in the data.
Fri Jan 24 15:44:17 +0000 2020FYI, reanalysis of #PXD014000 shows that the protocol for recovering unblocked cysteine-containing peptides in mammalian samples — J. R.Wiśniewski, et al., 🔗 — works as well as most protocols that use IAA/CAA blocking.
Tue Jan 14 15:29:52 +0000 2020@d_vivian That is a good suggestion (PXD004452). I particularly like the two 46 fraction MudPit runs with files labelled like "SH-SY5Y_REP1_46frac" & "SH-SY5Y_REP2_46frac". They are pretty consistent and were probably the best in terms of protein detection of the methods tried.
Mon Jan 13 20:51:24 +0000 2020@rbharathkumar91 @KentsisResearch @Ciencia_2017 A C. elegans dataset that maybe has a bit too much E. coli B for comfort is PXD004104 (🔗). One of the MudPit runs is given here, filtered to show the 1,641 E. coli protein ids 🔗
Thu Jan 09 17:59:10 +0000 2020Oddly, PXD015442 seems to be the first rat blood plasma data set made publicly available.
Sat Dec 28 18:51:44 +0000 2019If anyone is interested in checking an algorithm's ability to detect the normally rare arginine PTM that forms citrulline (deimidation), PXD012122 is a good data set to use.
Thu Dec 12 22:34:35 +0000 2019PXD013453: good chance to test out my "# of phosphosites nearing an asymptote" hypothesis
Thu Nov 28 17:38:05 +0000 2019Anyone who wants an ideal QE MS/MS data set to use for testing algorithms should give PXD007940 a try.
Thu Nov 28 17:22:05 +0000 2019#PXD007940: wirklich nette Daten! 👍
Mon Nov 18 13:43:40 +0000 2019DYNC1LI1:p, dynein cytoplasmic 1 light intermediate chain 1 (H. sapiens) 🔗 Midsized cytosolic protein; 2 phosphodomains; 1 SAAVs: M147T (maf=0.02); mature form 2-523 [30,585 x] 🔗

Wed Oct 16 16:41:24 +0000 2019@neely615 You can use CAA, but you should change your alkylation protocols and check the recovery of Cys-containing peptides. As it turns out, even the group that promoted the change isn't immune to having CAA-alkylation fail without noticing it (e.g., PXD010697)
Sat Oct 12 15:50:29 +0000 2019#PXD008600 is quite interesting. The composition of the cervical mucus plugs don't conform with my expectations of the constituents of a ball of mucus, but that may be just me. I would never have guessed that fibroblasts were involved in forming these little rascals.
Mon Oct 07 17:08:58 +0000 2019@Smith_Chem_Wisc @neely615 A good example is PXD010154, try "01087_C01_P010739_S00_N03_R1.raw", "01087_H01_P010739_S00_N08_R1.raw" and or "01087_F01_P010739_S00_N06_R1.raw".
Wed Oct 02 18:42:00 +0000 2019@KentsisResearch @Smith_Chem_Wisc @CompProteomics PXD010154 is also pretty deep and has a nice range of tissues, just be ready for lots of carbamylation.
Wed Oct 02 18:38:12 +0000 2019@KentsisResearch @Smith_Chem_Wisc @CompProteomics The raw files labelled "70frac_no_concatenation_rep1" or "293_REP1_46frac" in PXD004452 are pretty good for cell lines and "PROSTATE_Pt2_human_46frac" for tissue.
Sat Sep 28 02:12:05 +0000 2019A nearly ideal MHC II peptide length distribution, from an analysis of PXD014253 (q2133257c.raw), part of 🔗 🔗
Mon Sep 23 18:22:13 +0000 2019If you are interested in succinylation and how to detect it in proteomics, #PXD013173 from 🔗 is a good data set to try out.
Tue Jul 23 21:40:44 +0000 2019When someone does a really good job on MHC class 2 peptides (e.g., PXD013599), the results are fascinating.
Mon Jul 08 16:00:32 +0000 2019PXD006503 is pretty interesting data. It suggests that it would be valuable to have a lot more K-malonylation data to compare with what is known about K-ubiquitinylation.
Thu May 23 15:30:01 +0000 2019The dataset PXD010982 (🔗, 🔗, human hair) is a good example of how complicated it can be to distinguish between closely related proteins even with lots of sample & relatively few distinct gene products present.
Fri May 10 16:53:34 +0000 2019@TrostLab @KentsisResearch To be sure, I grabbed a run from PXD010621 & ranked the proteins by # of phospho-PSMs detected. Of the top 20 SRRM2 NCL TRIM28 HSP90AA1 SRRM1 TCOF1 MKI67 AKAP12 BCLAF1 ATRX MSH6 PCM1 RANBP2 ZC3H13 MAP1B TRP53BP1 NPM1 TJP2 NUCKS1 RIF1 all but 2 are nuclear.
Tue Apr 30 18:41:03 +0000 2019@HFazelinia @BrenesAlejandro Doll S, et al. Nat Commun. 2017 (PXD006675, 🔗) generated the best MT-ND6 ids in heart that I've seen. There is only 1 good tryptic peptide, but there are several semi-tryptic peptides that show up if you use "value-priced" Lys-C.
Fri Apr 05 20:23:34 +0000 2019#PXD010175 (Yang M, et al. 🔗) is really solid, thorough proteomics data. If you want to know what proteins are observable using TMT6 in human B cells, this one is about as good as it gets 🔗
Thu Mar 28 14:56:28 +0000 2019#PXD009040 is a good quality, very interesting data set (🔗). Has the high ZNF-domain protein concentration (~10% of ids) characteristic of sumoylation affinity capture studies.
Wed Mar 06 16:58:29 +0000 2019Is there any way to discourage people from making PRIDE depositions like 🔗 (PXD011839)? By placing all of the raw files in one very large zip file, it makes downloading the data unnecessarily difficult and error prone.
Wed Feb 27 13:50:21 +0000 2019@LM_Orre Good old fashioned mislabeling. The unique peptide MCS plot should have been labeled blue (lower curve) PXD010154 and the orange (upper curve) PXD006895.
Wed Feb 27 13:47:49 +0000 2019This plot is mislabeled. The blue (lower curve) should be PXD010154 and the orange (upper curve) should be PXD006895.
Wed Feb 20 15:34:24 +0000 2019Both studies are near the maximum number of unique proteins that can be discovered using their experimental protocols. The experimental procedure used in PXD006895 generated almost 60% more unique protein ids.
Wed Feb 20 15:28:36 +0000 2019This plot shows the unique protein detection performance of 2 recent large-scale proteomics experiments (PXD010154 🔗 & PXD006895 🔗). Both studies used about 1.3 million MS/MS spectra per sample. 🔗
Fri Feb 08 02:19:16 +0000 2019Always something unexpected: PXD011762 is the first data set I've seen where the tryptic peptides have their cysteines cysteinylated (the method says they should be IAA derivatized). It is common in MHC peptides, but not in tryptic digests.
Thu Jan 24 17:42:06 +0000 2019PXD009283: really good data ✓, interesting problem ✓, charismatic species ✓, 🔗
Tue Jan 15 21:48:38 +0000 2019#PXD003912 is one of the very (very) few data sets I have seen in which the experimental protocol results in complete recovery of cysteine-containing tryptic peptides. Good work 🔗
Thu Dec 06 18:33:06 +0000 2018#PXD009781 has genuinely good data, but the samples are really tricky things to model accurately. I would be interested to see what an "open" search engine would do with this sort of data
Wed Dec 05 19:00:13 +0000 2018#PXD007662 is a really well done: one of the best data sets I've seen. It shows how to design and execute an experiment for the reliable detection of SAVs.
Mon Nov 19 17:43:27 +0000 2018PXD011622 certainly lives up to its stated goal of providing a good sampling of N. meningitidis outer membrane proteins.
Thu Nov 15 13:49:20 +0000 2018PXD011271 is a really nice data set. Interesting technically and biologically. Good work Dr. Foster, et al.
Wed Aug 29 17:45:36 +0000 2018#PXD008914 does a surprisingly good job of isolating RNA-binding proteins from mESCs. If you are reanalyzing it, watch out for semi-tryptic & incompletely labelled peptides.
Tue Aug 28 21:17:16 +0000 2018@NCI_NCIP @theNCI @NCItreatment @NCICancerCtrl @NCI_CSSI The corresponding PRIDE proteomics data set PXD008832 still hasn't been released.
Mon Aug 27 15:15:27 +0000 2018#PXD010025 is a good example of how careful about your assumptions you must be when analyzing data from samples containing a lot of extracellular matrix.
Tue Aug 14 19:11:54 +0000 2018#PXD008478 & #PXD008473 are very good data sets that can be used by anyone interested in understanding how a larger-than-human vertebrate proteome (Oncorhynchus mykiss) affects statistics and spectra-peptide-protein matching process.
Fri Aug 10 01:15:31 +0000 2018#PXD007906 is one of the more interesting data sets I've seen in a while: lots of uncirculated plasma proteins exactly as they were when they were secreted from hepatocytes (HepG2 in this case).
Tue Aug 07 14:43:46 +0000 2018I'm still on the fence about #PXD008127. I've analyzed a sample of 36 runs and some are good sets of HLA type I (TI) peptides, but most are mixtures of TI peptides and endogenous plasma peptides - in some cases with no discernible TI peptides at all.
Wed Aug 01 14:31:38 +0000 2018If you want to study the effects of unanticipated non-tryptic cleavage on a conventional proteomics experiment, #PXD007234 is the perfect dataset to examine.
Mon Jul 30 00:59:08 +0000 2018#PXD008057 Pacific oyster data is really interesting: who knew they had so much collagen 1, with so many modifications. Would have been better for me without the TMT, though.
Sat Jul 21 14:38:21 +0000 2018#PXD005945 is much better and more interesting than I thought it would be. Worth taking a look at if you are interested in using effective protein separations for high sensitivity proteomics. Good data for course work assignments, too.
Wed Jun 06 14:20:06 +0000 2018#PXD007826 has some really nice chromatography and good reproducibility 🔗 A bit more carbamylation than I like to see, but not to a level that would adversely effect the analysis.
Tue May 22 13:18:45 +0000 2018#PXD002121 (from 2015, 🔗) still holds up as one of the truly great proteomics data sets
Sun May 20 16:55:17 +0000 2018# PXD008199 good data set if you are interested in mouse saliva
Sat May 19 21:02:22 +0000 2018#PXD008726 flamboyantly carbamylated, but great cysteine recovery.
Thu May 10 14:22:56 +0000 2018#PXD006109 is the last time I ever look at data from a Nat. Methods paper. What a waste of time!
Thu Apr 19 14:28:46 +0000 2018#PXD008840 has got to be the most interesting large scale cancer proteomics data set ever produced. A very preliminary analysis is available at 🔗

tweets = 289