Full-Length Vector Transcript
Built on target-capture enrichment combined with PacBio Iso-Seq full-length cDNA long-read sequencing, this service identifies vector-related full-length isoforms, host–vector fusion transcripts, alternative splicing, and 3′-end usage at the transcript level, providing single-molecule–level evidence for transgene expression integrity and insertion safety.
1. Background
Once a vector integrates into the host genome, it is transcribed into RNA and translated into protein (such as CAR). However, integration sites are open genomic environments where transcription does not always proceed cleanly, and several common transcript-level risks need to be considered.
First, host–vector fusion transcripts (VECTOR∷host). Transcription may read through the vector into a neighboring host gene, or read from a host gene into the vector, producing chimeric fusion RNAs that can both affect CAR expression and serve as transcript-level evidence of whether the insertion influences nearby gene transcription. Second, 3′-end anomalies caused by cryptic polyadenylation cause premature transcription termination and compromise transcript integrity. Third, alternative splicing produces multiple isoform versions from the same gene region, and it is necessary to confirm that the dominant version is indeed the functional transcript intended by design.
An isoform refers to a distinct RNA version generated from the same gene through different splicing or start/end sites. At the vector transcript level, the central questions are: what is the proportion of functional full-length transcripts, and are abnormal or fusion versions present? Short-read sequencing excels at quantification and per-base accuracy, but cannot independently reconstruct the complete transcript structure from 5′ to polyA, nor fusion junctions — and that is precisely the strength of long-read Iso-Seq.

Figure 1. Schematic of host–vector fusion transcripts and abnormal isoforms.
From a regulatory standpoint, transcript identification and expression-integrity characterization have a clear basis: ICH Q6B requires quality specifications and acceptance criteria including identity; the FDA guidance "Considerations for the Development of CAR T Cell Products" (Final, 2024) requires evaluation of identity, potency, and safety with attention to insertional mutagenesis risk; and China NMPA/CDE's "Technical Guideline for Pharmaceutical Research and Evaluation of Immune Cell Therapy Products (Trial)" (2022) sets explicit requirements for the pharmaceutical research and safety of CAR-T and other products. This service provides transcript-level characterization data for these objectives.
2. Technical Principle
This service is built on target-capture enrichment combined with PacBio Iso-Seq full-length cDNA long-read sequencing. Target capture uses probes designed against the customer-provided vector to enrich vector-related transcripts and is an included step of this service. Iso-Seq covers complete transcripts from 5′ to polyA in single reads, directly resolving isoform structure, fusion junctions, alternative splicing, and 3′-end usage without assembly. The core workflow is as follows:
(1) Sample receipt and QC
Receive cell pellets or extracted RNA; assess RNA concentration, purity, and integrity (such as RIN/DV200); and confirm compliance with library construction requirements.
(2) Full-length cDNA synthesis and custom probe capture enrichment
Full-length cDNA is synthesized, and probes customized to the customer's vector sequence are used to target-capture vector-related transcripts, improving the proportion of effective signal.
(3) PacBio Iso-Seq sequencing
Sequencing is performed on the PacBio HiFi platform to the target depth, with single-molecule long reads completely covering transcript structure.
(4) Bioinformatic analysis and result interpretation
Alignment is performed against the customer's vector sequence and the human reference genome as custom references. Full-length isoforms are identified, fusion junctions are detected, and alternative splicing and 3′-end usage are resolved. A transcript characterization report is then issued.

Figure 2. Workflow of target capture combined with PacBio Iso-Seq.
3. Technical Features and Advantages
(1) Single-molecule full-length reads with no assembly required
A single read fully covers transcripts from 5′ to polyA, directly revealing isoform structure and fusion junctions without assembly and avoiding short-read assembly ambiguity.
(2) Localization of fusion transcripts
Identifies and localizes host–vector fusion-transcript junctions, providing key transcript-level evidence on whether insertion affects neighboring gene expression.
(3) Comprehensive view of splicing and 3′-end usage
Resolves alternative splicing events and 3′-end usage, identifying premature termination caused by cryptic polyadenylation and abnormal isoforms.
(4) Custom alignment to the client's vector
Read-by-read alignment is performed against the customer-provided full vector sequence as a custom reference, with conclusions that are quantitative, locus-resolved, and comparable across runs.
(5) Complementary to DNA-level evidence
Can be jointly interpreted with the Vector Integration Site and Integrant Structure Analysis to build a complete safety argument that combines "where the integration is" with "what is transcribed."
4. Applications
Characterization of transgene expression integrity: Confirm that the dominant full-length isoform matches the design, ruling out truncated or aberrant versions.
Transcript-level evidence on insertion safety: Identify host–vector fusion transcripts and cross-corroborate with DNA-level integration-site evidence.
Verification of vector design optimization: Verify transcript-level improvements after vector design or process optimization.
IND/BLA submission data support: Provide transcript identification and expression-integrity evidence for regulatory submissions.
5. Report and Deliverables
The report provides quantitative, locus-resolved, comparable transcript-level evidence. Core contents include:
·Structure and abundance of vector-related full-length isoforms, including the proportion of functional full-length isoforms and a catalog of abnormal versions.
·Identification and junction localization of host–vector fusion transcripts.
·Alternative splicing event and 3′-end usage analysis.
·Joint interpretation with the integration-site or integrant-structure analysis (optional, provided when combined interpretation is selected).
·Data deliverables: Complete analysis report (PDF), isoform and fusion-event result files, and raw sequencing data.
6. Service Workflow
Service Step | Description |
Project consultation and study design | Target-capture and sequencing plan designed according to vector type, study stage, and regulatory objectives. |
Sample receipt and QC | Assessment of RNA concentration, purity, and integrity. |
Full-length cDNA synthesis and custom capture | Probes customized to the customer's vector are used to enrich vector-related transcripts. |
High-throughput sequencing | Sequencing on the PacBio HiFi platform to the target depth. |
Bioinformatic analysis | Using the customer's vector sequence as a custom reference; isoforms, fusions, splicing, and 3′-end usage are resolved. |
Report delivery and technical support | Complete analysis report (PDF), result files, and raw data, plus follow-up technical consultation. |
* Standard turnaround: 40–45 business days.
7. Sample Requirements
Item | Submission Requirement |
Sample type | Cell pellets or extracted RNA. |
Recommended input | Total RNA ≥1 μg recommended (refer to the latest Sample Submission Guide; submit sufficient overhead above the minimum input). |
Concentration and quality | RNA concentration ≥50 ng/μL recommended; OD260/280 ≈ 1.8–2.1; RIN ≥7 or DV200 ≥70%; no significant degradation. |
Storage and shipping | Store at −80 °C; ship on dry ice with continuous cold-chain. |
* The latest Sample Submission Guide takes precedence. This service is not applicable to severely degraded samples. Please schedule and confirm the study plan before sample submission.
8. Technical Specifications
Parameter | Description |
Sequencing platform | PacBio HiFi (long-read single-molecule sequencing). |
Read strategy | Single reads covering complete transcripts (from 5′ to polyA). |
Enrichment strategy | Target capture with probes customized to the customer's vector (included step of this service). |
Applicable sample | Cell pellets or extracted RNA. |
Alignment reference | Customer-provided full vector sequence and the human reference genome as custom references. |
Detection capability | Full-length isoform structure, host–vector fusion transcripts, alternative splicing, and 3′-end usage. |
Sequencing depth | Set according to target sensitivity (detection of low-abundance isoforms and fusion events scales with depth). |
Method status | IND: fit-for-purpose method qualification; BLA: full validation per ICH Q2(R2). |
Species supported | Primarily human cells; other species accommodated using the customer's vector sequence. |
9. References
[1] ICH. Q6B: Specifications: Test Procedures and Acceptance Criteria for Biotechnological/Biological Products. Current Step 4 version, 10 March 1999.
[2] ICH. Q2(R2): Validation of Analytical Procedures. Step 4 version, adopted 1 November 2023.
[3] U.S. Food and Drug Administration (FDA), Center for Biologics Evaluation and Research (CBER). Considerations for the Development of Chimeric Antigen Receptor (CAR) T Cell Products. Final, January 2024. (Docket No. FDA-2021-D-0404)
[4] Center for Drug Evaluation, National Medical Products Administration of China (NMPA-CDE). Technical Guideline for Pharmaceutical Research and Evaluation of Immune Cell Therapy Products (Trial) [in Chinese]. Notice No. 11 of 2022, issued 2022.