README file for Chamaecrista fasciculata (Cf) transcript assemblies Steven Cannon, steven.cannon@ars.usda.gov or scannon@iastate.edu April, 2009 DATA SOURCE: These files are the result of a NSF SGER-funded project, PI Susan Singer, and a collaboration with Jeff Doyle, Greg May, Steven Cannon, Sonja Maki, and Dan Illut. Please contact Susan Singer (ssinger@carleton.edu) with questions about use of the pre-publication data. FILES: INTERMEDIATE FILE Chafa_454ctg.ffn (renamed from 454-assembly-consensi.seq) Result of contigging 454 Titanium sequences. letters sequences min_len max_len ave_len 124,438,993 367,802 43 2,019 338.33 INTERMEDIATE FILE Chafa0.5_cDNA_2008-12.ffn (renamed from Cf_asmbly2_08-12-28.ffn) Result of extending 454 contigs with Solexa sequence. This is likely to be significantly under-clustered. letters sequences min_len max_len ave_len 34,289,821 60,015 43 5,740 571.35 FILE Chafa1.0_cDNA_2009-03.ffn Result of additional collapsing of similar contigs in the Chafa0.5 file above. letters sequences min_len max_len ave_len 32,890,172 54,903 43 5,758 599.06 FILE Chafa1.1clean_cDNA_2009-04.ffn Result of selecting regions homologous to soybean predicted peptides, with corrections for probable frame shifts, then applying cap3 to further collapse contigs. letters sequences min_len max_len ave_len 10,086,547 21,781 48 3,376 463.09 METHODS: Plants from Chamaecrista fasciculata ecotype MN98 were grown for tissue. Tissues were collected from 14 conditions, including root, shoot, nodule, and inflorescence tissues at several time points. Tissue was prepared for mRNA library production. Libraries were sequenced using two methods: with Roche 454 Titanium on pooled RNA from all 12 tissues to generate long reads to serve as alignment templates; and with the Illumina Genome Analyzer to generate high sequence coverage and transcript counts (for use in analyses to be published separately). For the Roche 454 Titanium sequencing, cDNA was synthesized with the following adaptors, and subsequently sheared during 454 library prep. 5'end-AAGCAGTGGTATCAACGCAGAGTGGCCATTACGGCCGGG-cDNA- AAAAAAAAAAGAAAAAAAAACAAAACATGTCGGCCGCCTCGGTCTCTA-3'end The total number of reads in the bulk run was 950,227 with an average read-length of 344.69bp and a total number of bases of 327,532,623. For short-read generation (ÒSolexaÓ sequence), Whole transcriptome shotgun (WTS) libraries were generated, and were sequenced on an Illumina Genome Analyzer to a read length of 46 base pairs. The average number of reads per library was 9.428 million, and total reads were 132 million, and 6 billion DNA bases Assembly proceeded by contigging the 454 reads, followed by several rounds of contig extensions. The 454 reads were contigged using a custom perl pipeline and Blat (Kent, 2002) to bin homologous sequences, then cap3 (Hwang and Madan, 1999) to align sequences in contigs. Solexa reads were filtered for quality as follows. Low-scoring right tails were trimmed from any bases with phred < 4 (after conversion to phred scores with a custom perl script). Additionally, the following low-scoring sequences were removed: sequences consisting of greater than 66% of poly-A -T -G -C; and sequences with fewer than 30 high quality bases. Contig generation and extension with Solexa sequence proceeded using VCAKE (Jeck et al., 2007), adding sequences in four batches (to avoid computer memory limitations in the 32 bit perl implementation, as VCAKE is written in perl). Each round used VCAKE parameters -k 42 (i.e. use 42 nucleotides from each sequence); -n 21 (percent overlap required; default is 18). The first VCAKE round extended from the 454 contigs, and next two round extended from the preceding contigs. After the third round, contigs were collapsed using cap3 (Hwang and Madan, 1999), using default parameters. In the final round of Solexa extensions, contigs and singletons from cap3 were extended with the remaining batch of Solexa reads. This set of preliminary contigs (Chafa0.5_cDNA_2008-12.ffn) consisted of 60,015 contigs, with average length 571 nt. Evaluation of the initial contigs by homology comparisons of the sequences with themselves indicated under-contigging, with highly similar sequences probably remaining from allelic variants (from heterozygous source DNA), splice variants, and minor sequencing variants. Therefore, preparatory to alignments and tree construction, we used two additional rounds of contigging with cap3, followed by conceptual translation using exonerate (version 2.2; Slater and Birney, 2005) to align each Cf sequence in-frame to the most similar soybean peptide sequence from the Glyma1.01 annotation. We used the exonerate alignments and custom perl scripts to remove single-base insertions responsible for probable frame shifts, and a final run of cap3 to re-collapse in-frame sequences resulting from the exonerate comparisons. The resulting file, Chafa1.1clean_cDNA_2009-04.ffn, has 21,781 sequences, with average length 463 nt. Of these sequences, 90.7% are without stop codons in frame 0. The remainder may either be ORFs in other frames, or may contain frame shifts.