Kraków / Remote · open to new opportunities

Marcin RapaczSenior Frontend Developer / AI Engineer

Frontend developer, six years on a single product. Since 2026, mostly applications with an AI layer, and more and more often I step outside the front end too: I write the backend and take projects all the way to a running server.

LinkedIn ↗

01Experience

05.2026 — present

Fullstack Developer / AI Engineer

Nodeic — own company (B2B contract) · client: AI recruitment platform (NDA)

Building a recruitment platform with an AI-powered candidate search engine. I built the front end from scratch; the backend, the vector database and the infrastructure were already running, so I work inside existing code — the search engine, the data model, migrations and re-processing the whole base after changes. The same client also got an internal contact-sourcing tool.

  • Built a production web application from scratch (Nuxt 4 / Vue 3, SSR, BFF): Vertical Slice Architecture on Nuxt Layers, a design system of my own, PL/EN i18n, consent-gated analytics.
  • Developing a semantic candidate search engine (Google Gemini, e5-large embeddings, Qdrant) over a corpus of a few thousand CVs: strategy-pattern rebuild, hard filters, pipeline versioning and a full re-extraction run in production. I wrote the second version of the engine test-first — the scoring engine is fully covered, and edge-case model responses are exercised through a fake LLM, because writing a test is easier than forcing a model into a specific wrong answer.
  • Improved the data the search runs on: a canonical profession category derived during CV extraction, skill entries reduced to a shared form, and a rebuild of the matching logic with Polish inflection handling. see the case study ↓
  • Established a search evaluation methodology (LLM-as-judge, ranking correlations, a 50-posting benchmark built from real-traffic logs) and based decisions on it: data fixes raised ranking agreement with the independent evaluation from 2% to 52%, my cross-encoder reranking R&D ended in a documented decision not to ship with the conclusion that the next leap needed LLM reranking, and a second engine version built in isolation from production — with a light LLM reranker built on that conclusion — wins the benchmark 22:2 (ordering agreement 70% vs 35%) and serves production traffic — the post-deployment balance, computed on 1,214 rows of real search logs, showed four classes of defect gone from the list the client sees, among them a filter cutting more than half the pool in 9% of queries and 6% of positions scoring below the quality threshold. see the case study ↓
  • GDPR in the AI layer: fail-closed anonymization of surnames in LLM-generated text (handling Polish declension). Found and fixed a stored XSS in the CV preview myself: the file came back with the MIME type stored at upload and opened as a blob inheriting the panel's origin — the server now forces the content type, adds nosniff and rejects anything that is not a PDF.
  • For the same client I built the second version of a B2B lead-sourcing tool (Python, Pydantic AI / Gemini, FastAPI, PostgreSQL, Docker), rewritten from scratch: a pipeline from job postings, through employer profiles and model enrichment with grounding, to three independent sources verifying the phone number — the company's website, the GUS state register, Google Places — in “free before paid” order. Every fact carries its origin and status, the model runs in a single typed call forbidden to guess, and its output feeds a deterministic target filter. About 1.3k companies in the base, 83% of in-target companies with a confirmed phone. see the case study ↓

03.2020 — 05.2026

Letter of recommendation ↗

Frontend Developer

Printbox Sp. z o.o.

A product company with its own photo-product personalization editors and hundreds of clients worldwide. Six years on the e-commerce team: from React redesigns of the editors' mobile versions, through building a headless Shopware frontend from demo to production, to a multi-tenant architecture and a framework migration on a live product.

  • Co-designed a multi-tenant architecture for three client tiers and owned its full implementation — from panel-only configuration to full redesigns (Turborepo, Nuxt Layers, UI/logic component split, a feature configuration plugin): new stores are built as layers over a shared core, with no forks and no regressions for other clients.
  • Owned the migration from Nuxt 2 to Nuxt 3 on a live product — others joined when they had free cycles: moving the shared demo and the new build system; along the way I moved configuration into Redis and started reading it at runtime, which ended data drift between instances and application rebuilds after every admin-panel change. see the case study ↓
  • Built the front end and its integrations for the rollout of the platform's most important client, a US retail chain: a full redesign, SSO and a new payment gateway, in collaboration with the backend team. For the pilot of a global electronics manufacturer entering the photo market I built most of the front end. see the case study ↓
  • Embedded the company's flagship photo-product editor (React) inside the Nuxt stores — reconciling two routers, browser navigation support, runtime-configured languages; built a Cloudflare Workers prerender for its client-side-rendered purchase paths (SEO, saving a few hundred USD/month). see the case study ↓
  • Rebuilt the build and deployment system: one Docker image configured via environment variables instead of a build per client, slimmed from 1.4 GB to 55 MB — updating every store takes minutes instead of hours, regardless of client count.
  • Integrated payments (Stripe, Adyen) and shipping providers (InPost, GLS) in the critical purchase path, inside legacy Shopware plugins.
  • Daily with the backend, and with product before a task: consultations on what could be done better, cheaper or faster, planning and estimates for tasks and redesigns. Often I was the one digging into new solutions worth pushing the product towards — not always the originator, but the person who researched the topic to the end. Informal mentoring: people came to me for the architecture and the product — not five-minute chats, but sessions of tens of minutes; on top of that, demos of recent changes for the team and documentation wherever it was worth writing.

02Own projects

2026 — present

View demo ↗

Fullstack Developer / AI Engineer

StoryForge — own SaaS product · built after hours

Creator and sole author of a SaaS platform for book authors: an editor with version history, RAG chat over the author's own book, style analysis and a credit system.

  • Built an editor with version history and autosave, separating the working save, the point in history and the index the AI reads from — the interface never pretends the model knows newer text than it actually does. see the case study ↓
  • Designed and shipped a PDF book-import agent (FastAPI + LangGraph + Gemini, behind a proxy in Strapi): chapters are detected in code — the PDF outline or a heading scan with numbering-continuity checks — and the model only gets a say when several conventions compete, picking from a closed list; when no split can be established the import stops instead of inventing chapters. Instead of a progress bar the author reads the agent's narration (25 event codes in PL/EN), and a failed import refunds credits through transactional sessions in Strapi. 569 offline tests with a hard block on live model calls.
  • Chat and passage analysis over the author's own book (Next.js, FastAPI, Qdrant): six analysis modes, four of which deliberately never query the vector database, retrieval by the vector of the passage rather than the question, hard filters, and a plain “nothing found” message. see the case study ↓
  • Wrote the style analysis without a model — my own heuristics for sentence boundaries, dialogue and repetition within a window, with versioned metrics. No inference cost and the same result on every run.
  • AI jobs go through a queue in Postgres and credits through a transactional ledger that refunds when the model stream dies before the first token, so a failure on the AI side never blocks saving text.
  • I run the whole infrastructure myself: services behind Traefik, Docker, GitLab CI/CD pipelines and scripted provisioning from a bare server, with E2E smoke tests that make no paid AI calls.

Frontend Developer

Folstar — company website, mikroperforacja.pl · freelance commission, working directly with the client

Company website of a family-run foil packaging manufacturer and sweets wholesaler near Kraków, rebuilt from scratch on Next.js 16 and React 19. A commission run directly with the client, no intermediaries: scope, copy and SEO, launch and maintenance — everything a small company needs from one person.

  • Solutions sized to the client: content in code instead of the previous version's headless CMS — with a dozen or so products that means fewer services to maintain and no dependency on an external vendor, and every content change goes through the same path as code.
  • The rest is the ordinary craft of a company website, done properly: View Transitions in React 19 that respect the system's reduced-motion setting, an accessible product dialog, Analytics and the Google map only after consent, local SEO (LocalBusiness in JSON-LD), a Docker image behind Traefik deployed from GitLab CI to my own VPS.

03Skills

Frontend

  • TypeScript
  • Vue / Nuxt
  • React / Next.js
  • View Transitions (React 19)
  • Tailwind CSS
  • Pinia
  • TanStack Query
  • VueUse
  • Zustand
  • shadcn/ui + Radix
  • i18n (pluralisation, localised routes)
  • Zod (validating backend responses)
  • Figma (pixel-perfect handoff, designs read via MCP)

Backend

  • Node.js
  • Strapi
  • Python / FastAPI
  • SQLAlchemy
  • PostgreSQL
  • Redis

AI & Data

  • RAG
  • Vector databases (Qdrant)
  • Embeddings
  • Evaluation (LLM-as-judge)
  • Gemini API
  • Vercel AI SDK (streaming)
  • LangChain / LangGraph
  • LangSmith
  • Pydantic AI
  • Tools for the model (function calling)

Architecture & DevOps

  • Vertical Slice
  • BFF
  • Docker
  • Traefik
  • GitLab CI/CD
  • VPS

Quality & security

  • Vitest
  • pytest
  • Playwright
  • Core Web Vitals
  • Sentry
  • GDPR / data anonymisation
  • GA4 + consent mode

Process

  • AI-Driven Development (Claude Code: TDD, subagents, I read and correct every change myself)
  • Agile / Scrum
  • Code Review

Languages

  • Polish — native
  • English — B2

04Selected case studies

Short engineering stories — the problem, the decision, and the measured outcome.

Recruitment platform (NDA)

A benchmark settled which search engine is better

Measurement came first: an independent AI judge that never sees the system's own scoring, and a benchmark built from the logs of real searches, with expectations written down by hand. The first round showed the problem was the data, not the model — fixing the skill entries raised ranking agreement with the independent evaluation from 2% to 52%, and the cross-encoder reranking R&D ended in a documented “don't ship” — with the conclusion that the next leap needed LLM reranking. Then I went further: a second version of the engine built in isolation from production, the judge calibrated like a measuring instrument (repeatability, a practical ceiling of ≈ 80%), ranking weights derived from correlations computed per posting, a light LLM reranker ordering the shortlist with no right to reject — that conclusion made real — and 440+ offline tests on a faked model that returns raw text.

Across the 50 benchmark postings the new version puts its list closer to the AI report in 22 cases, the baseline in 2; ordering agreement 70% vs 35%, the top pick right 47% vs 36% of the time. The new version serves production traffic — the post-deployment balance, computed on 1,214 log rows, showed four classes of defect gone from the list the client sees, and the mean Spearman correlation with the judge on real searches holds at 0.72 (108 searches, as of September 2026).

Read the full case study

Recruitment platform (NDA)

Vectors couldn't tell professions apart

A query for an accountant returned financial controllers and analysts — their competencies are similar, and the job title itself didn't stand out strongly enough in the vector to cut neighbouring roles off. The effect was the opposite of what was intended: some actual accountants dropped out of the top ten in favour of people from adjacent professions. I built a canonical profession category derived during CV extraction and wired it into search alongside vector similarity. Rather than a yes/no filter it is a directional map of partial credit between professions: a financial controller can still surface for an accountant query, but with a lower weight, and a threshold cuts off whatever is too far away. That threshold only started working once the whole base had been re-processed — while some CVs had no category, the “unknown passes” escape hatch was masking the problem.

The role stopped being guessed from text similarity: the category answers who someone is, and the vectors answer what the category doesn't describe — experience, industry and the detail of their skills.

Internal tool for a client in HR (NDA)

No fact from the model reaches a person without confirmation

The first version was an agent: the model drove the run at seven points, and an LLM gate before sending, in a blind test on 17 contacts, flagged every defective one and rejected none of the good ones. The second version, rewritten from scratch, shifts the weight from the model's decisions to verifiable facts. The system builds its own base of companies from job postings and employer profiles, and a model with search access adds company size, phone and website — in a single typed call, forbidden to guess, at temperature zero. An answer that fails schema validation never reaches the database; a low-confidence fact is never created at all. Every fact has a source and a status: unverified, confirmed or rejected. Confirmation has to come from an independent source, and the order of the checks was set by cost: first the company's website and the GUS state register, because they are free, and only then paid Google Places — and only when two free queries point at the same company. The spending fuse lives in the database and survives a restart, the register's rate limit is enforced in the client, and every model call has its cost recorded. When an audit showed company websites planting junk numbers, only full numbers were confirmed from then on, and 51 earlier ones were moved to rejected.

About 1.3k companies in the base: 83% of in-target companies have a phone confirmed by an independent source (684 of 825), and 459 companies dropped out as off-target before the verification stage. Over 300 offline tests with network blocked and a fake model.

StoryForge

Not every question needs the vector database

Analysis of a selected passage has six modes, from sentence rhythm to consistency of the fictional world, and the mode decides two things at once: what the model should do, and whether the vector database is queried at all. Four of the six never touch it. Where retrieval does make sense, I search with the vector of the passage itself rather than the question: text similar to a paragraph is where the same characters and facts live, which is exactly where a contradiction hides. The question joins the vector only in the free-form mode, where there is no telling in advance what the author is after. The chunk containing the analysed passage is dropped from the results so the model never compares the text with itself. The passage and its context window go into the system instruction, not into the conversation thread.

How much the model knows follows from the chosen mode rather than from chance. And when a mode needs the rest of the book and nothing is found, the model says so outright instead of judging consistency from the passage alone.

StoryForge

Autosave, save and the AI index are three different states

The author writes, autosave runs in the background every two seconds, but the model answers from an index that only catches up after a deliberate save. I separated the three states instead of pretending everything is in sync: the “AI has the current version” indicator appears only when the server confirms it and there are no unsaved changes on screen. Style statistics are computed by the backend from what is in the database, so before the panel opens I flush whatever autosave hasn’t sent yet — and if that fails, the panel doesn’t open at all.

The author always knows what they are working with: style analysis either shows numbers from the current text or doesn't open at all, and the chat says outright when it is answering from a stale index.

Printbox

A React editor inside a Nuxt store

The flagship photo-product editor had its own purchase paths and its own router (React), while the stores ran on Nuxt with languages configured at runtime. Embedding it meant reconciling two routers: browser back and forward buttons, correct remounts, and no hardcoded paths anywhere. The second problem followed from the first — the editor's purchase paths render on the client, so search engine crawlers were served an empty page. I built them a prerender on Cloudflare Workers that returns ready HTML from cache.

The company's flagship product sells inside every store on the platform as if it were native to it, and its purchase paths are visible to search engines — the prerender also replaced a paid external service.

Printbox

Nuxt 2 to Nuxt 3 on a live product

The platform moved to a new framework without pausing sales. I owned moving the shared demo every store is built on, and the new build system — thanks to the shared core most stores came across with the demo, and only clients with their own code changes needed separate work. Along the way I moved configuration into Redis and started reading it at runtime: before that some settings were baked in at build time, so a change in the admin panel forced an application rebuild, and across many instances the data could drift between them.

Fully customised clients were deliberately left until last, so the hardest cases went through a path that was already proven. I don't remember a single deployment we had to roll back — the new deployment path made it possible within minutes, but we never needed it.

Printbox

Rollouts for the platform's key clients

The platform's most important client, a US retail chain, moved to the new stack together with a full redesign. On the front-end side I led that rollout: moving the store onto the new codebase, sign-in through the client's own system (SSO), and wiring a new payment gateway into the critical purchase path — both integrations in cooperation with the backend team, on a store that kept selling throughout. For the pilot of a global electronics manufacturer testing entry into the photo-product market I built most of the front end: a full redesign matching the client's visual identity, and the front-end integrations.

The retail chain moved onto the new platform, and the pilot looked as though it had been built from scratch for that brand — while both stores ran on the same codebase as every other rollout.

05How I work

Working on a team

On the e-commerce team we kept a full rhythm: standups, planning, demos, retros. Changes went through a ticket, a branch and a merge request, so someone else looked at every one before it was merged. Shared code with other frontend developers, conflicts to resolve, and day-to-day work with backend engineers, testers and customer success.

Owning the task

I take a task whole: from understanding the problem to a working release. Usually the goal is enough — the order of steps and the solutions along the way are mine to settle. Before that I pin down the details, and sometimes suggest a different approach that gets the same result for less.

Code written with AI

I write with Claude Code, and every feature goes through the same cycle: analysis and a plan, implementation in TDD with subagents for partial tasks, tests and benchmarks, deployment through CI/CD. The code that comes out of it I read and correct the way I would someone else's in review, before it reaches the repository. That pace produces more code than it used to, so all the more reason to record in commits and notes why something looks the way it does.

Tests where failure is silent

The worst bugs don't crash the app. They shift numbers or reorder results and look perfectly plausible. Those are the places I try to cover most densely with tests. The rest is checked too, in proportion to what can go wrong.

Design system and Figma

I start from colour, spacing and radius tokens kept in one place, and collect recurring layouts into shared components. Views go in pixel-perfect from Figma, and the first draft of a component I pull straight from the design through the Figma MCP server rather than retyping values by hand.

Performance isn't a separate phase

Core Web Vitals, bundle size and behaviour on a slow connection get checked alongside the feature I'm building. On the data side it's usually the same moves: collapse several queries into one, cache what doesn't change by the minute, and return only the fields the view actually renders.

06Training & education

Contact

Kraków or remote. I work in Polish and in English. Write if you'd like to talk about working together, or to ask about any of the projects.

I use Google Analytics to understand how this site is used and what to improve. Analytics scripts load only after you give consent. Privacy policy

Get in touch

Leave a message — I usually reply within one business day.