7 Open-Source AI Projects, Ranked by How Easily an Accountant Can Contribute

Last Saturday night, watching the rain pour down on Seattle through the window, I had about twelve GitHub tabs open. One was related to LLM fine-tuning, one was an MLOps pipeline, and another was a data labeling tool. Every time I switched tabs, I'd read the contributing.md, scan through the issue list, and try to figure out what I could actually contribute. Most of the projects? No place for me. I can't optimize an inference engine in C++, and I'm not writing custom PyTorch layers.

But after repeating this for several months, I started to notice a pattern. What open-source AI projects truly lack isn't just code. Documentation cleanup, cost analysis, benchmark data validation, license review, governance documentation. There are clearly areas where someone who's good with spreadsheets and obsessive about making the numbers add up is needed. So I put together a list of the projects I actually looked into. The ranking is based on "how meaningfully can I, as an accountant, contribute."

Starting from the lowest barrier to contribution

7th. llama.cpp. Honestly, this was mostly just spectating. It's C/C++-centric, and the PR reviews require serious technical depth. That said, spreadsheet skills come in handy for organizing benchmark results or creating hardware-specific performance comparison tables. I once posted a neatly formatted table of performance data on an issue, and a maintainer replied with thanks. That was all it was, but not a bad start.

6th. Hugging Face's datasets library. There's room to contribute in areas like writing dataset cards or validating metadata. The work involves organizing data sources, licenses, sizes, and formats according to a standard template, which is structurally similar to formatting audit reports. The onboarding docs are written with Python developers in mind, though, so it took me a while to get to my first PR.

A spreadsheet showing organized benchmark data with colored headers

5th. MLflow. It's an experiment tracking and model registry tool, and the documentation is vast and constantly needs updating. My area was catching cases where code examples in tutorials didn't match the latest API, and writing feature proposals related to cost tracking. When I submitted a feature request for tracking cloud costs on a per-experiment basis, it got a pretty positive response.

4th. Community projects related to the Open LLM Leaderboard. The core work involves collecting and comparing model benchmark scores, and data integrity validation is always backlogged. Tracking the submission history of models with suspiciously high scores, or separating data from before and after changes in evaluation criteria. It's well suited for someone accustomed to making every single number line up.

3rd. The Kubernetes ecosystem, particularly cost optimization tools. Rather than k8s itself, there are open-source tools in the surrounding ecosystem that analyze cluster costs. This is where understanding cloud billing structures gives you an advantage. You can review the logic for calculating cost efficiency relative to resource usage, or contribute to price table updates. I never expected my habit of poring over the AWS billing dashboard every month to be useful here.

2nd. AI ethics and governance framework projects. This includes things like model card authoring standards and AI system audit checklists. Reading regulatory documents like the NIST AI Framework and converting them into practical checklists is remarkably similar to writing compliance reports. The same meticulousness you use in internal control audits carries right over. The community also tends to be open to non-developer contributions, so the psychological barrier was low.

A desk with financial documents next to a laptop showing a code repository

1st. Financial transparency documentation for nonprofit AI projects. Specifically, this involves reviewing financial reports published by open-source foundations and AI-related nonprofit organizations, then summarizing them in a format the community can understand. Extracting R&D spending ratios from annual reports, categorizing funding sources, and visualizing trends in expenditure items. This is exactly my day job. It's structurally identical to working with Form 990s for nonprofit clients during tax season.

What I learned from onboarding

A common frustration was projects that say "non-code contributions are welcome" but provide no actual guidance on where to start. You open issues tagged "good first issue," and they're all code changes. The more clearly a project documented its non-developer contribution paths, the less time it took me to actually open a PR.

One more thing. Learning Git properly as I started contributing to open source was an unexpected bonus. Creating branches, rebasing, resolving conflicts. Considering I used to version-control in Excel with "final_final_reallyFinal.xlsx," I've come a long way.

Even as I write this, I have three new tabs open. One is a stock screener, one is a Kubernetes forum, and one is an AI ethics paper I haven't read yet. I'm not closing any of them. I plan to go through them one by one over coffee tomorrow morning.

Comments