
CodeT5
Open source code understanding and generation models from Salesforce Research used for translation summarization and synthesis across many programming languages.
Overview
CodeT5 and CodeT5 plus are encoder decoder models tuned for code tasks like generation translation explanation and summarization. Released by Salesforce Research with permissive assets these models appear in research baselines and applied systems where open weights are preferred. Developers use the checkpoints to bootstrap assistants fine tune on domain repositories or evaluate RAG pipelines against coding tasks.
The repo includes scripts data links and usage examples for common frameworks so labs can reproduce results and extend experiments. Because the models are open they support offline experiments and custom governance which is useful for sensitive environments and education settings that cannot send code to third party clouds.
Key features
- Open weights and examples for research and applied prototypes
- Supports generation summarization translation and explanation
- Encoder decoder design with variants for different sizes
- Reference scripts datasets and evaluation guidance
- Strong baselines on public coding benchmarks
- Compatible with popular deep learning frameworks
- Community issues and PRs improve docs and utilities
- Permissive use in offline or private settings subject to license
Best for
- Bootstrap code assistants without external API reliance
- Translate between languages or frameworks for migrations
- Summarize long source files or PRs for reviewers
- Label functions and generate docstrings for clarity
- Build evaluation harnesses for coding tasks and RAG
- Teach students about program synthesis with open weights
- Run ablations to test prompting finetuning and data
- Prototype domain adapters for internal stacks safely
Capabilities
Synthesis and Docstrings
Produce functions comments and docstrings to speed documentation and onboarding in open environments.
Language to Language
Convert snippets across languages or frameworks to aid migrations and code exploration.
Long Files
Create concise overviews of modules so reviewers and students grasp intent quickly.
Fine Tuning
Start from open checkpoints and tune on domain repositories to align with internal patterns.
Frequently Asked Questions
How does pricing start?
The models and code are available at no cost under the project license for research and applied experimentation.
Do you host an API?
No this is an open source release; serve locally or on your own cloud.
What datasets are used?
The repo links papers and datasets used to train and evaluate variants with details for reproduction.
Can I use it commercially?
Follow the repository license and referenced datasets’ terms before shipping products.
Is GPU required?
Inference can run on consumer GPUs for smaller checkpoints while larger variants need more VRAM.



