refactoring, small fixes, readme added
AREADME.md
@@ -0,0 +1,90 @@
# VidArchive
A self-hosted web front end for [yt-dlp](https://github.com/yt-dlp/yt-dlp). It
queues downloads, keeps the results in a browsable library with playback,
metadata, subtitles and comments, and re-runs saved subscriptions on a schedule.
Go standard library plus chi, SQLite (pure-Go driver), and server-rendered
templates. No JavaScript, no build step for the front end.
## Requirements
- `yt-dlp` on `PATH` (or set `VIDARCHIVE_YTDLP_PATH`)
- `ffmpeg` and `ffprobe` for thumbnails, subtitle conversion, and media probing
- Go 1.26+ to build from source
Missing tools are reported at startup and on `/healthz`; the app still starts.
## Run
With the container image:
```sh
docker compose up -d # or: podman-compose up -d
```
The compose file mounts `./data` and publishes port 8080. The image installs
yt-dlp at build time, so rebuild to update it:
```sh
docker compose build --pull --no-cache vidarchive
```
From source:
```sh
go build -o vidarchive ./cmd/vidarchive
./vidarchive
```
Then open <http://localhost:8080>.
## Configuration
All configuration is environment variables. Templates and static assets are
embedded in the binary.
| Variable | Default | Purpose |
| --- | --- | --- |
| `VIDARCHIVE_PORT` | `8080` | Listen port |
| `VIDARCHIVE_DATA_DIR` | `./data` | Root for everything below |
| `VIDARCHIVE_DB_PATH` | `<data>/vidarchive.db` | SQLite database |
| `VIDARCHIVE_LIBRARY_DIR` | `<data>/library` | Imported media |
| `VIDARCHIVE_TEMP_DIR` | `<data>/temp` | Download scratch space |
| `VIDARCHIVE_YTDLP_PATH` | `yt-dlp` | yt-dlp binary |
| `VIDARCHIVE_FFMPEG_PATH` | `ffmpeg` | ffmpeg binary |
| `VIDARCHIVE_FFPROBE_PATH` | `ffprobe` | ffprobe binary |
| `VIDARCHIVE_WORKERS` | `2` | Concurrent downloads (minimum 1) |
| `VIDARCHIVE_SCHEDULER_INTERVAL` | `60` | Seconds between subscription checks |
| `VIDARCHIVE_BASE_URL` | — | External URL; an `https://` value enables HSTS |
## Concepts
**Presets** collect the yt-dlp options for a download: format selection, audio
extraction, subtitle and thumbnail embedding, info-JSON and comment collection,
plus free-form custom flags. One preset can be the default. A download may
override the format and add its own flags.
**Queue.** A download row is claimed by a worker, which runs yt-dlp into a
scratch directory and then imports each finished item into the library. Live
output is streamed into memory and flushed to the row periodically, so the
detail page shows progress. Stopping the server leaves in-flight downloads
`downloading`; they are re-queued on the next start.
**Library items** are directories holding one or more media files, a
`.vidarchive-item.toml` marker, yt-dlp's `info.json`, thumbnails, and an
optional `subtitles/` directory. The marker is the source of truth for the
item's name, source URL, video id, and per-file durations, so listing pages
never have to run ffprobe. Directories without a marker are shown as folders,
which makes the library browsable as a tree.
**Subscriptions** re-download a URL on a cron schedule into a directory they
own. Three refresh modes:
- `overwrite` — replace the existing copy of each item in place
- `skip` — keep a yt-dlp download archive and fetch only new entries
- `metadata` — refresh metadata for known items, download only genuinely new ones
With *prune removed* enabled, items no longer present upstream are deleted
locally. Pruning is skipped when the source cannot be enumerated, so a network
error cannot empty the directory.
Mcompose.yml
@@ -14,7 +14,9 @@ services:
- VIDARCHIVE_PORT=8080
- VIDARCHIVE_DATA_DIR=/data
- VIDARCHIVE_WORKERS=2
- VIDARCHIVE_REFRESH_INTERVAL=5
# How often the subscription scheduler looks for due runs, in seconds. The
# page auto-refresh interval is a UI setting, not an environment variable.
- VIDARCHIVE_SCHEDULER_INTERVAL=60
# Uncomment to run behind HTTPS proxy:
# - VIDARCHIVE_BASE_URL=https://vidarchive.example.com
# Uncomment to use a custom yt-dlp path:
Minternal/handler/handler.go
@@ -8,11 +8,9 @@ import (
"net/http"
"net/url"
"path/filepath"
"strconv"
"strings"
"github.com/gabriel-vasile/mimetype"
"github.com/go-chi/chi/v5"
"vidarchive"
"vidarchive/internal/config"
@@ -245,612 +243,3 @@ func (h *Handler) settingsOrDefault() *models.Settings {
}
return settings
}
func (h *Handler) Library(w http.ResponseWriter, r *http.Request) {
path := r.URL.Query().Get("path")
filter := r.URL.Query().Get("filter")
sortBy := sortFromRequest(w, r, "library_sort", "date")
items, folders, err := h.librarySvc.GetAll(path, sortBy, filter)
if err != nil {
h.serverError(w, r, "list library", err)
return
}
settings := h.settingsOrDefault()
h.renderWithRequest(w, r, "library", PageData{
Title: "Library",
ActiveTab: "library",
AutoRefresh: settings.AutoRefreshLibrary,
RefreshSec: settings.RefreshInterval,
Data: struct {
Items []*models.LibraryItem
Folders []string
Path string
SortBy string
Filter string
}{
Items: items,
Folders: folders,
Path: path,
SortBy: sortBy,
Filter: filter,
},
})
}
func normalizeRelPath(r *http.Request) string {
// chi gives the raw, still-encoded wildcard. Decode it as a URL path, where
// '+' is a literal plus (only query strings treat '+' as space) — so an item
// directory named "a+b" round-trips correctly.
relPath := chi.URLParam(r, "*")
relPath = strings.Trim(relPath, "/")
if decoded, err := url.PathUnescape(relPath); err == nil {
relPath = decoded
} else {
log.Printf("normalizeRelPath: undecodable path %q: %v", relPath, err)
}
return relPath
}
func (h *Handler) LibraryItem(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
if r.Method == "POST" && strings.HasSuffix(relPath, "/delete") {
relPath = strings.TrimSuffix(relPath, "/delete")
h.deleteMedia(relPath, w, r)
return
}
h.libraryDetail(relPath, w, r)
}
const commentPreviewLimit = 50
func (h *Handler) LibraryComments(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
h.libraryComments(relPath, w, r)
}
func (h *Handler) libraryComments(relPath string, w http.ResponseWriter, r *http.Request) {
item, err := h.librarySvc.GetByRelPath(relPath)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
comments, _, err := h.librarySvc.GetEngagement(relPath)
if err != nil {
h.serverError(w, r, "load comments", err)
return
}
h.renderWithRequest(w, r, "library_comments", PageData{
Title: "Comments - " + item.Name,
ActiveTab: "library",
Data: struct {
Item *models.LibraryItem
Comments []models.Comment
}{Item: item, Comments: comments},
})
}
func (h *Handler) libraryDetail(relPath string, w http.ResponseWriter, r *http.Request) {
item, err := h.librarySvc.GetByRelPath(relPath)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
selectedFilename := r.URL.Query().Get("file")
if selectedFilename == "" && len(item.MediaFiles) > 0 {
selectedFilename = item.MediaFiles[0].Filename
}
meta, err := h.librarySvc.GetMetadata(relPath, selectedFilename)
if err != nil {
log.Printf("libraryDetail %q: metadata unavailable: %v", relPath, err)
}
subtitles, err := h.librarySvc.GetSubtitles(relPath)
if err != nil {
log.Printf("libraryDetail %q: subtitles unavailable: %v", relPath, err)
}
comments, heatmap, err := h.librarySvc.GetEngagement(relPath)
if err != nil {
log.Printf("libraryDetail %q: engagement unavailable: %v", relPath, err)
}
previewComments := comments
if len(previewComments) > commentPreviewLimit {
previewComments = previewComments[:commentPreviewLimit]
}
h.renderWithRequest(w, r, "library_detail", PageData{
Title: item.Name,
ActiveTab: "library",
Data: struct {
Item *models.LibraryItem
SelectedFilename string
Metadata *service.MediaMetadata
Subtitles []models.SubtitleTrack
Comments []models.Comment
CommentTotal int
Heatmap []models.HeatmapSegment
}{
Item: item,
SelectedFilename: selectedFilename,
Metadata: meta,
Subtitles: subtitles,
Comments: previewComments,
CommentTotal: len(comments),
Heatmap: heatmap,
},
})
}
func (h *Handler) ServeMediaItem(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
if strings.HasSuffix(relPath, "/thumbnail") {
h.serveThumbnail(strings.TrimSuffix(relPath, "/thumbnail"), w, r)
return
}
// Subtitle tracks are addressed as <item>/subtitles/<lang>, with the language
// as a trailing path segment (see LibraryService.GetSubtitles).
if i := strings.LastIndex(relPath, "/subtitles/"); i >= 0 {
item := relPath[:i]
lang := relPath[i+len("/subtitles/"):]
h.serveSubtitles(item, lang, w, r)
return
}
h.serveMedia(relPath, w, r)
}
func (h *Handler) serveMedia(relPath string, w http.ResponseWriter, r *http.Request) {
filename := r.URL.Query().Get("file")
if filename == "" {
http.Error(w, "Missing file", http.StatusBadRequest)
return
}
mediaPath, err := h.librarySvc.GetMediaFile(relPath, filename)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
http.ServeFile(w, r, mediaPath)
}
func (h *Handler) serveThumbnail(relPath string, w http.ResponseWriter, r *http.Request) {
filename := r.URL.Query().Get("file")
if thumb, ok := h.librarySvc.ThumbnailForFile(relPath, filename); ok {
http.ServeFile(w, r, thumb)
return
}
// Fall back to an icon, matched to the requested file's type (or the item's
// primary file when no specific file was requested).
icon := "video-icon.svg"
if item, err := h.librarySvc.GetByRelPath(relPath); err == nil {
if isAudioFile(item, filename) {
icon = "audio-icon.svg"
}
}
data, err := vidarchive.StaticFS.ReadFile("web/static/icons/" + icon)
if err != nil {
http.Error(w, "icon not found", http.StatusInternalServerError)
return
}
w.Header().Set("Content-Type", "image/svg+xml")
w.Write(data)
}
// isAudioFile reports whether the named file (or, if unnamed, the first media
// file) of an item is audio.
func isAudioFile(item *models.LibraryItem, filename string) bool {
if filename != "" {
for _, mf := range item.MediaFiles {
if mf.Filename == filename {
return mf.IsAudio
}
}
return false
}
return len(item.MediaFiles) > 0 && item.MediaFiles[0].IsAudio
}
func (h *Handler) serveSubtitles(relPath, lang string, w http.ResponseWriter, r *http.Request) {
if lang == "" {
http.Error(w, "Missing language", http.StatusBadRequest)
return
}
if strings.Contains(lang, "/") || strings.Contains(lang, "..") || strings.Contains(lang, "\\") {
http.Error(w, "Invalid language", http.StatusBadRequest)
return
}
subtitlePath, err := h.librarySvc.GetSubtitlePath(relPath, lang)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
w.Header().Set("Content-Type", "text/vtt")
http.ServeFile(w, r, subtitlePath)
}
func (h *Handler) deleteMedia(relPath string, w http.ResponseWriter, r *http.Request) {
if err := h.librarySvc.Delete(relPath); err != nil {
redirectWithError(w, r, "/library", "Couldn't delete this item.", err)
return
}
redirectWithSuccess(w, r, "/library", "Item deleted.")
}
func (h *Handler) Downloads(w http.ResponseWriter, r *http.Request) {
status := r.URL.Query().Get("status")
sortBy := sortFromRequest(w, r, "queue_sort", "date")
downloads, err := h.downloadSvc.GetAll(status, sortBy)
if err != nil {
h.serverError(w, r, "list downloads", err)
return
}
settings := h.settingsOrDefault()
h.renderWithRequest(w, r, "queue", PageData{
Title: "Queue",
ActiveTab: "queue",
AutoRefresh: settings.AutoRefreshDownloads,
RefreshSec: settings.RefreshInterval,
Data: struct {
Items []*models.Download
Status string
SortBy string
}{
Items: downloads,
Status: status,
SortBy: sortBy,
},
})
}
func (h *Handler) DownloadDetail(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
download, err := h.downloadSvc.GetByID(id)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
h.renderWithRequest(w, r, "queue_detail", PageData{
Title: "Queue Details",
ActiveTab: "queue",
Data: download,
})
}
func (h *Handler) CreateDownload(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
url := r.FormValue("url")
if url == "" {
flashError(w, "A URL is required to start a download.")
http.Redirect(w, r, "/download", http.StatusSeeOther)
return
}
var presetID *int64
if pid := r.FormValue("preset_id"); pid != "" {
id, err := strconv.ParseInt(pid, 10, 64)
if err == nil {
presetID = &id
}
}
formatOverride := r.FormValue("format_override")
customFlags := r.FormValue("custom_flags")
outputDir := r.FormValue("output_dir")
download, err := h.downloadSvc.Create(url, presetID, formatOverride, customFlags, outputDir)
if err != nil {
redirectWithError(w, r, "/download", "Couldn't queue this download.", err)
return
}
h.workerPool.Submit(download)
redirectWithSuccess(w, r, "/queue", "Download queued.")
}
func (h *Handler) DeleteDownload(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := h.downloadSvc.Delete(id); err != nil {
redirectWithError(w, r, "/queue", "Couldn't remove this download.", err)
return
}
redirectWithSuccess(w, r, "/queue", "Download removed.")
}
func (h *Handler) ClearAllDownloads(w http.ResponseWriter, r *http.Request) {
if err := h.downloadSvc.DeleteAll(); err != nil {
redirectWithError(w, r, "/queue", "Couldn't clear the queue.", err)
return
}
redirectWithSuccess(w, r, "/queue", "Queue cleared.")
}
func (h *Handler) Settings(w http.ResponseWriter, r *http.Request) {
presets, err := h.presetSvc.GetAll()
if err != nil {
h.serverError(w, r, "list presets", err)
return
}
settings, err := h.settingsSvc.GetAll()
if err != nil {
h.serverError(w, r, "load settings", err)
return
}
h.renderWithRequest(w, r, "settings", PageData{
Title: "Settings",
ActiveTab: "settings",
Data: struct {
Presets []*models.Preset
Settings *models.Settings
}{
Presets: presets,
Settings: settings,
},
})
}
func (h *Handler) CreatePreset(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
preset := &models.Preset{}
if err := applyPresetForm(preset, r); err != nil {
redirectWithError(w, r, "/settings", err.Error(), nil)
return
}
if err := h.presetSvc.Create(preset); err != nil {
redirectWithError(w, r, "/settings", "Couldn't create this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset created.")
}
// applyPresetForm copies the preset form fields onto p and validates them. It is
// shared by create and update so the two can't drift apart as fields are added.
func applyPresetForm(p *models.Preset, r *http.Request) error {
name := strings.TrimSpace(r.FormValue("name"))
if name == "" {
return fmt.Errorf("A preset needs a name.")
}
// Mirrors the radio options on the settings form; empty means "unspecified"
// and BuildArgs applies its own default.
formatMode := r.FormValue("format_mode")
switch formatMode {
case "", "default", "preset", "custom":
default:
return fmt.Errorf("Unknown format mode %q.", formatMode)
}
p.Name = name
p.Description = r.FormValue("description")
p.FormatMode = formatMode
p.Format = r.FormValue("format")
p.Quality = r.FormValue("quality")
p.CustomFormat = r.FormValue("custom_format")
p.AudioFormat = r.FormValue("audio_format")
p.SubLangs = r.FormValue("sub_langs")
p.CustomFlags = r.FormValue("custom_flags")
p.IsDefault = r.FormValue("is_default") == "1"
p.ExtractAudio = r.FormValue("extract_audio") == "1"
p.EmbedSubs = r.FormValue("embed_subs") == "1"
p.EmbedThumbnail = r.FormValue("embed_thumbnail") == "1"
p.EmbedMetadata = r.FormValue("embed_metadata") == "1"
p.WriteInfoJSON = r.FormValue("write_info_json") == "1"
p.WriteComments = r.FormValue("write_comments") == "1"
p.CommentSort = strings.TrimSpace(r.FormValue("comment_sort"))
p.CommentExtractorArgs = strings.TrimSpace(r.FormValue("comment_extractor_args"))
p.MaxComments = 0
if raw := strings.TrimSpace(r.FormValue("max_comments")); raw != "" {
maxComments, err := strconv.Atoi(raw)
if err != nil || maxComments < 0 {
return fmt.Errorf("Max comments must be a non-negative number.")
}
p.MaxComments = maxComments
}
// Comments are stored in the info JSON sidecar. Keep the dependent options
// consistent even when a client submits the form without JavaScript.
if !p.WriteInfoJSON || !p.WriteComments {
p.WriteComments = false
p.CommentSort = ""
p.MaxComments = 0
p.CommentExtractorArgs = ""
}
return nil
}
func (h *Handler) UpdatePreset(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
preset, err := h.presetSvc.GetByID(id)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
if err := applyPresetForm(preset, r); err != nil {
redirectWithError(w, r, "/settings", err.Error(), nil)
return
}
if err := h.presetSvc.Update(preset); err != nil {
redirectWithError(w, r, "/settings", "Couldn't update this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset updated.")
}
func (h *Handler) DeletePreset(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := h.presetSvc.Delete(id); err != nil {
redirectWithError(w, r, "/settings", "Couldn't delete this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset deleted.")
}
func (h *Handler) UpdateSettings(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
var firstErr error
record := func(err error) {
if err != nil && firstErr == nil {
firstErr = err
}
}
if interval := r.FormValue("refresh_interval"); interval != "" {
record(h.settingsSvc.SetRefreshInterval(interval))
}
record(h.settingsSvc.SetAutoRefreshLibrary(r.FormValue("auto_refresh_library") == "1"))
record(h.settingsSvc.SetAutoRefreshDownloads(r.FormValue("auto_refresh_downloads") == "1"))
record(h.settingsSvc.SetCookies(r.FormValue("cookies")))
if firstErr != nil {
log.Printf("UpdateSettings: %v", firstErr)
flashError(w, "Some settings couldn't be saved: "+firstErr.Error())
} else {
flashSuccess(w, "Settings saved.")
}
http.Redirect(w, r, "/settings", http.StatusSeeOther)
}
func (h *Handler) DownloadForm(w http.ResponseWriter, r *http.Request) {
presets, err := h.presetSvc.GetAll()
if err != nil {
h.serverError(w, r, "list presets", err)
return
}
// A missing default preset is normal (the user may not have set one); the
// template handles a nil DefaultPreset, so this isn't surfaced as an error.
defaultPreset, _ := h.presetSvc.GetDefault()
url := r.URL.Query().Get("url")
var formats []*models.FormatInfo
var flash *Flash
if r.URL.Query().Get("list_formats") == "1" && url != "" {
var err error
if formats, err = h.downloadSvc.ListFormats(url); err != nil {
log.Printf("DownloadForm: listing formats for %q failed: %v", url, err)
flash = &Flash{Kind: "error", Message: "Couldn't list formats: " + err.Error()}
}
}
h.renderWithRequest(w, r, "download_form", PageData{
Title: "Download",
ActiveTab: "download",
Flash: flash,
Data: struct {
Presets []*models.Preset
DefaultPreset *models.Preset
URL string
FormatOverride string
CustomFlags string
OutputDir string
Formats []*models.FormatInfo
ShowFormats bool
}{
Presets: presets,
DefaultPreset: defaultPreset,
URL: url,
FormatOverride: r.URL.Query().Get("format_override"),
CustomFlags: r.URL.Query().Get("custom_flags"),
OutputDir: r.URL.Query().Get("output_dir"),
Formats: formats,
ShowFormats: r.URL.Query().Get("list_formats") == "1",
},
})
}
func (h *Handler) GetPresetFlags(w http.ResponseWriter, r *http.Request) {
idStr := r.URL.Query().Get("id")
id, err := strconv.ParseInt(idStr, 10, 64)
if err != nil {
http.Error(w, "Invalid ID", http.StatusBadRequest)
return
}
preset, err := h.presetSvc.GetByID(id)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
flags := h.presetSvc.EffectiveFlags(preset, "", "")
w.Header().Set("Content-Type", "text/plain")
w.Write([]byte(flags))
}
func (h *Handler) Theme(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
theme := r.FormValue("theme")
if theme == "" {
theme = "auto"
}
setCookie(w, "theme", theme)
referer := r.Header.Get("Referer")
if referer == "" {
referer = "/"
}
http.Redirect(w, r, referer, http.StatusSeeOther)
}
Ainternal/handler/helpers_test.go
@@ -0,0 +1,148 @@
package handler
import (
"context"
"net/http"
"net/http/httptest"
"net/url"
"testing"
"github.com/go-chi/chi/v5"
)
// A flash must be shown exactly once: reading it also expires the cookie.
func TestConsumeFlashClearsCookie(t *testing.T) {
w := httptest.NewRecorder()
setFlash(w, "success", "Preset created.")
cookie := w.Result().Cookies()[0]
r := httptest.NewRequest("GET", "/settings", nil)
r.AddCookie(cookie)
next := httptest.NewRecorder()
flash := consumeFlash(next, r)
if flash == nil {
t.Fatal("no flash read back")
}
if flash.Kind != "success" || flash.Message != "Preset created." {
t.Errorf("flash = %+v", flash)
}
cleared := next.Result().Cookies()
if len(cleared) != 1 || cleared[0].MaxAge >= 0 {
t.Errorf("flash cookie was not expired: %+v", cleared)
}
}
// The cookie is client-editable, so an arbitrary kind must not reach the
// banner's class name.
func TestConsumeFlashClampsKind(t *testing.T) {
cases := map[string]string{
"success": "success",
"error": "error",
"evil\" onload=alert1": "error",
}
for kind, want := range cases {
r := httptest.NewRequest("GET", "/", nil)
r.AddCookie(&http.Cookie{Name: flashCookie, Value: url.QueryEscape(kind + "|hello")})
flash := consumeFlash(httptest.NewRecorder(), r)
if flash == nil {
t.Fatalf("kind %q: no flash read back", kind)
}
if flash.Kind != want {
t.Errorf("kind %q became %q, want %q", kind, flash.Kind, want)
}
}
}
func TestConsumeFlashIgnoresMalformedValues(t *testing.T) {
for _, value := range []string{"", "no-separator", "%zz"} {
r := httptest.NewRequest("GET", "/", nil)
r.AddCookie(&http.Cookie{Name: flashCookie, Value: value})
if flash := consumeFlash(httptest.NewRecorder(), r); flash != nil {
t.Errorf("value %q produced a flash: %+v", value, flash)
}
}
// No cookie at all.
if flash := consumeFlash(httptest.NewRecorder(), httptest.NewRequest("GET", "/", nil)); flash != nil {
t.Errorf("missing cookie produced a flash: %+v", flash)
}
}
// An explicit ?sort= is remembered, and a request without one restores the last
// choice — the auto-refresh reloads these pages without the query string.
func TestSortFromRequest(t *testing.T) {
w := httptest.NewRecorder()
r := httptest.NewRequest("GET", "/library?sort=title", nil)
if got := sortFromRequest(w, r, "library_sort", "date"); got != "title" {
t.Fatalf("explicit sort = %q, want title", got)
}
cookies := w.Result().Cookies()
if len(cookies) != 1 || cookies[0].Name != "library_sort" || cookies[0].Value != "title" {
t.Fatalf("sort was not remembered: %+v", cookies)
}
next := httptest.NewRequest("GET", "/library", nil)
next.AddCookie(cookies[0])
if got := sortFromRequest(httptest.NewRecorder(), next, "library_sort", "date"); got != "title" {
t.Errorf("restored sort = %q, want title", got)
}
// With neither query nor cookie, the caller's default applies.
bare := httptest.NewRequest("GET", "/library", nil)
if got := sortFromRequest(httptest.NewRecorder(), bare, "library_sort", "date"); got != "date" {
t.Errorf("default sort = %q, want date", got)
}
}
// The wildcard arrives still percent-encoded. '+' is a literal plus in a path,
// so an item directory named "a+b" must round-trip.
func TestNormalizeRelPath(t *testing.T) {
cases := map[string]string{
"folder/item": "folder/item",
"/folder/item/": "folder/item",
"a%2Bb": "a+b",
"a+b": "a+b",
"spaced%20name": "spaced name",
"%E6%97%A5%E6%9C%AC": "日本",
}
for raw, want := range cases {
r := httptest.NewRequest("GET", "/library/item/"+raw, nil)
ctx := chi.NewRouteContext()
ctx.URLParams.Add("*", raw)
r = r.WithContext(context.WithValue(r.Context(), chi.RouteCtxKey, ctx))
if got := normalizeRelPath(r); got != want {
t.Errorf("normalizeRelPath(%q) = %q, want %q", raw, got, want)
}
}
}
func TestParseIDRejectsNonNumeric(t *testing.T) {
r := httptest.NewRequest("POST", "/queue/abc/delete", nil)
ctx := chi.NewRouteContext()
ctx.URLParams.Add("id", "abc")
r = r.WithContext(context.WithValue(r.Context(), chi.RouteCtxKey, ctx))
w := httptest.NewRecorder()
if _, ok := parseID(w, r); ok {
t.Error("parseID accepted a non-numeric id")
}
if w.Code != http.StatusBadRequest {
t.Errorf("status = %d, want 400", w.Code)
}
}
func TestFormatDuration(t *testing.T) {
cases := map[int]string{
-1: "--:--",
0: "--:--",
59: "0:59",
61: "1:01",
3661: "1:01:01",
}
for seconds, want := range cases {
if got := formatDuration(seconds); got != want {
t.Errorf("formatDuration(%d) = %q, want %q", seconds, got, want)
}
}
}
Ainternal/handler/library.go
@@ -0,0 +1,255 @@
package handler
import (
"log"
"net/http"
"net/url"
"strings"
"github.com/go-chi/chi/v5"
"vidarchive"
"vidarchive/internal/models"
"vidarchive/internal/service"
)
func (h *Handler) Library(w http.ResponseWriter, r *http.Request) {
path := r.URL.Query().Get("path")
filter := r.URL.Query().Get("filter")
sortBy := sortFromRequest(w, r, "library_sort", "date")
items, folders, err := h.librarySvc.GetAll(path, sortBy, filter)
if err != nil {
h.serverError(w, r, "list library", err)
return
}
settings := h.settingsOrDefault()
h.renderWithRequest(w, r, "library", PageData{
Title: "Library",
ActiveTab: "library",
AutoRefresh: settings.AutoRefreshLibrary,
RefreshSec: settings.RefreshInterval,
Data: struct {
Items []*models.LibraryItem
Folders []string
Path string
SortBy string
Filter string
}{
Items: items,
Folders: folders,
Path: path,
SortBy: sortBy,
Filter: filter,
},
})
}
func normalizeRelPath(r *http.Request) string {
// chi gives the raw, still-encoded wildcard. Decode it as a URL path, where
// '+' is a literal plus (only query strings treat '+' as space) — so an item
// directory named "a+b" round-trips correctly.
relPath := chi.URLParam(r, "*")
relPath = strings.Trim(relPath, "/")
if decoded, err := url.PathUnescape(relPath); err == nil {
relPath = decoded
} else {
log.Printf("normalizeRelPath: undecodable path %q: %v", relPath, err)
}
return relPath
}
func (h *Handler) LibraryItem(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
if r.Method == "POST" && strings.HasSuffix(relPath, "/delete") {
relPath = strings.TrimSuffix(relPath, "/delete")
h.deleteMedia(relPath, w, r)
return
}
h.libraryDetail(relPath, w, r)
}
const commentPreviewLimit = 50
func (h *Handler) LibraryComments(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
h.libraryComments(relPath, w, r)
}
func (h *Handler) libraryComments(relPath string, w http.ResponseWriter, r *http.Request) {
item, err := h.librarySvc.GetByRelPath(relPath)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
comments, _, err := h.librarySvc.GetEngagement(relPath)
if err != nil {
h.serverError(w, r, "load comments", err)
return
}
h.renderWithRequest(w, r, "library_comments", PageData{
Title: "Comments - " + item.Name,
ActiveTab: "library",
Data: struct {
Item *models.LibraryItem
Comments []models.Comment
}{Item: item, Comments: comments},
})
}
func (h *Handler) libraryDetail(relPath string, w http.ResponseWriter, r *http.Request) {
item, err := h.librarySvc.GetByRelPath(relPath)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
selectedFilename := r.URL.Query().Get("file")
if selectedFilename == "" && len(item.MediaFiles) > 0 {
selectedFilename = item.MediaFiles[0].Filename
}
meta, err := h.librarySvc.GetMetadata(relPath, selectedFilename)
if err != nil {
log.Printf("libraryDetail %q: metadata unavailable: %v", relPath, err)
}
subtitles, err := h.librarySvc.GetSubtitles(relPath)
if err != nil {
log.Printf("libraryDetail %q: subtitles unavailable: %v", relPath, err)
}
comments, heatmap, err := h.librarySvc.GetEngagement(relPath)
if err != nil {
log.Printf("libraryDetail %q: engagement unavailable: %v", relPath, err)
}
previewComments := comments
if len(previewComments) > commentPreviewLimit {
previewComments = previewComments[:commentPreviewLimit]
}
h.renderWithRequest(w, r, "library_detail", PageData{
Title: item.Name,
ActiveTab: "library",
Data: struct {
Item *models.LibraryItem
SelectedFilename string
Metadata *service.MediaMetadata
Subtitles []models.SubtitleTrack
Comments []models.Comment
CommentTotal int
Heatmap []models.HeatmapSegment
}{
Item: item,
SelectedFilename: selectedFilename,
Metadata: meta,
Subtitles: subtitles,
Comments: previewComments,
CommentTotal: len(comments),
Heatmap: heatmap,
},
})
}
func (h *Handler) ServeMediaItem(w http.ResponseWriter, r *http.Request) {
relPath := normalizeRelPath(r)
if strings.HasSuffix(relPath, "/thumbnail") {
h.serveThumbnail(strings.TrimSuffix(relPath, "/thumbnail"), w, r)
return
}
// Subtitle tracks are addressed as <item>/subtitles/<lang>, with the language
// as a trailing path segment (see LibraryService.GetSubtitles).
if i := strings.LastIndex(relPath, "/subtitles/"); i >= 0 {
item := relPath[:i]
lang := relPath[i+len("/subtitles/"):]
h.serveSubtitles(item, lang, w, r)
return
}
h.serveMedia(relPath, w, r)
}
func (h *Handler) serveMedia(relPath string, w http.ResponseWriter, r *http.Request) {
filename := r.URL.Query().Get("file")
if filename == "" {
http.Error(w, "Missing file", http.StatusBadRequest)
return
}
mediaPath, err := h.librarySvc.GetMediaFile(relPath, filename)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
http.ServeFile(w, r, mediaPath)
}
func (h *Handler) serveThumbnail(relPath string, w http.ResponseWriter, r *http.Request) {
filename := r.URL.Query().Get("file")
if thumb, ok := h.librarySvc.ThumbnailForFile(relPath, filename); ok {
http.ServeFile(w, r, thumb)
return
}
// Fall back to an icon, matched to the requested file's type (or the item's
// primary file when no specific file was requested).
icon := "video-icon.svg"
if item, err := h.librarySvc.GetByRelPath(relPath); err == nil {
if isAudioFile(item, filename) {
icon = "audio-icon.svg"
}
}
data, err := vidarchive.StaticFS.ReadFile("web/static/icons/" + icon)
if err != nil {
http.Error(w, "icon not found", http.StatusInternalServerError)
return
}
w.Header().Set("Content-Type", "image/svg+xml")
w.Write(data)
}
// isAudioFile reports whether the named file (or, if unnamed, the first media
// file) of an item is audio.
func isAudioFile(item *models.LibraryItem, filename string) bool {
if filename != "" {
for _, mf := range item.MediaFiles {
if mf.Filename == filename {
return mf.IsAudio
}
}
return false
}
return len(item.MediaFiles) > 0 && item.MediaFiles[0].IsAudio
}
func (h *Handler) serveSubtitles(relPath, lang string, w http.ResponseWriter, r *http.Request) {
if lang == "" {
http.Error(w, "Missing language", http.StatusBadRequest)
return
}
if strings.Contains(lang, "/") || strings.Contains(lang, "..") || strings.Contains(lang, "\\") {
http.Error(w, "Invalid language", http.StatusBadRequest)
return
}
subtitlePath, err := h.librarySvc.GetSubtitlePath(relPath, lang)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
w.Header().Set("Content-Type", "text/vtt")
http.ServeFile(w, r, subtitlePath)
}
func (h *Handler) deleteMedia(relPath string, w http.ResponseWriter, r *http.Request) {
if err := h.librarySvc.Delete(relPath); err != nil {
redirectWithError(w, r, "/library", "Couldn't delete this item.", err)
return
}
redirectWithSuccess(w, r, "/library", "Item deleted.")
}
Ainternal/handler/queue.go
@@ -0,0 +1,164 @@
package handler
import (
"log"
"net/http"
"strconv"
"vidarchive/internal/models"
)
func (h *Handler) Downloads(w http.ResponseWriter, r *http.Request) {
status := r.URL.Query().Get("status")
sortBy := sortFromRequest(w, r, "queue_sort", "date")
downloads, err := h.downloadSvc.GetAll(status, sortBy)
if err != nil {
h.serverError(w, r, "list downloads", err)
return
}
settings := h.settingsOrDefault()
h.renderWithRequest(w, r, "queue", PageData{
Title: "Queue",
ActiveTab: "queue",
AutoRefresh: settings.AutoRefreshDownloads,
RefreshSec: settings.RefreshInterval,
Data: struct {
Items []*models.Download
Status string
SortBy string
}{
Items: downloads,
Status: status,
SortBy: sortBy,
},
})
}
func (h *Handler) DownloadDetail(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
download, err := h.downloadSvc.GetByID(id)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
h.renderWithRequest(w, r, "queue_detail", PageData{
Title: "Queue Details",
ActiveTab: "queue",
Data: download,
})
}
func (h *Handler) CreateDownload(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
url := r.FormValue("url")
if url == "" {
flashError(w, "A URL is required to start a download.")
http.Redirect(w, r, "/download", http.StatusSeeOther)
return
}
var presetID *int64
if pid := r.FormValue("preset_id"); pid != "" {
id, err := strconv.ParseInt(pid, 10, 64)
if err == nil {
presetID = &id
}
}
formatOverride := r.FormValue("format_override")
customFlags := r.FormValue("custom_flags")
outputDir := r.FormValue("output_dir")
download, err := h.downloadSvc.Create(url, presetID, formatOverride, customFlags, outputDir)
if err != nil {
redirectWithError(w, r, "/download", "Couldn't queue this download.", err)
return
}
h.workerPool.Submit(download)
redirectWithSuccess(w, r, "/queue", "Download queued.")
}
func (h *Handler) DeleteDownload(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := h.downloadSvc.Delete(id); err != nil {
redirectWithError(w, r, "/queue", "Couldn't remove this download.", err)
return
}
redirectWithSuccess(w, r, "/queue", "Download removed.")
}
func (h *Handler) ClearAllDownloads(w http.ResponseWriter, r *http.Request) {
if err := h.downloadSvc.DeleteAll(); err != nil {
redirectWithError(w, r, "/queue", "Couldn't clear the queue.", err)
return
}
redirectWithSuccess(w, r, "/queue", "Queue cleared.")
}
func (h *Handler) DownloadForm(w http.ResponseWriter, r *http.Request) {
presets, err := h.presetSvc.GetAll()
if err != nil {
h.serverError(w, r, "list presets", err)
return
}
// A missing default preset is normal (the user may not have set one); the
// template handles a nil DefaultPreset, so this isn't surfaced as an error.
defaultPreset, _ := h.presetSvc.GetDefault()
url := r.URL.Query().Get("url")
var formats []*models.FormatInfo
var flash *Flash
if r.URL.Query().Get("list_formats") == "1" && url != "" {
var err error
if formats, err = h.downloadSvc.ListFormats(url); err != nil {
log.Printf("DownloadForm: listing formats for %q failed: %v", url, err)
flash = &Flash{Kind: "error", Message: "Couldn't list formats: " + err.Error()}
}
}
h.renderWithRequest(w, r, "download_form", PageData{
Title: "Download",
ActiveTab: "download",
Flash: flash,
Data: struct {
Presets []*models.Preset
DefaultPreset *models.Preset
URL string
FormatOverride string
CustomFlags string
OutputDir string
Formats []*models.FormatInfo
ShowFormats bool
}{
Presets: presets,
DefaultPreset: defaultPreset,
URL: url,
FormatOverride: r.URL.Query().Get("format_override"),
CustomFlags: r.URL.Query().Get("custom_flags"),
OutputDir: r.URL.Query().Get("output_dir"),
Formats: formats,
ShowFormats: r.URL.Query().Get("list_formats") == "1",
},
})
}
Ainternal/handler/settings.go
@@ -0,0 +1,205 @@
package handler
import (
"fmt"
"log"
"net/http"
"strconv"
"strings"
"vidarchive/internal/models"
)
func (h *Handler) Settings(w http.ResponseWriter, r *http.Request) {
presets, err := h.presetSvc.GetAll()
if err != nil {
h.serverError(w, r, "list presets", err)
return
}
settings, err := h.settingsSvc.GetAll()
if err != nil {
h.serverError(w, r, "load settings", err)
return
}
h.renderWithRequest(w, r, "settings", PageData{
Title: "Settings",
ActiveTab: "settings",
Data: struct {
Presets []*models.Preset
Settings *models.Settings
}{
Presets: presets,
Settings: settings,
},
})
}
func (h *Handler) CreatePreset(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
preset := &models.Preset{}
if err := applyPresetForm(preset, r); err != nil {
redirectWithError(w, r, "/settings", err.Error(), nil)
return
}
if err := h.presetSvc.Create(preset); err != nil {
redirectWithError(w, r, "/settings", "Couldn't create this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset created.")
}
// applyPresetForm copies the preset form fields onto p and validates them. It is
// shared by create and update so the two can't drift apart as fields are added.
func applyPresetForm(p *models.Preset, r *http.Request) error {
name := strings.TrimSpace(r.FormValue("name"))
if name == "" {
return fmt.Errorf("A preset needs a name.")
}
// Mirrors the radio options on the settings form; empty means "unspecified"
// and BuildArgs applies its own default.
formatMode := r.FormValue("format_mode")
switch formatMode {
case "", "default", "preset", "custom":
default:
return fmt.Errorf("Unknown format mode %q.", formatMode)
}
p.Name = name
p.Description = r.FormValue("description")
p.FormatMode = formatMode
p.Format = r.FormValue("format")
p.Quality = r.FormValue("quality")
p.CustomFormat = r.FormValue("custom_format")
p.AudioFormat = r.FormValue("audio_format")
p.SubLangs = r.FormValue("sub_langs")
p.CustomFlags = r.FormValue("custom_flags")
p.IsDefault = r.FormValue("is_default") == "1"
p.ExtractAudio = r.FormValue("extract_audio") == "1"
p.EmbedSubs = r.FormValue("embed_subs") == "1"
p.EmbedThumbnail = r.FormValue("embed_thumbnail") == "1"
p.EmbedMetadata = r.FormValue("embed_metadata") == "1"
p.WriteInfoJSON = r.FormValue("write_info_json") == "1"
p.WriteComments = r.FormValue("write_comments") == "1"
p.CommentSort = strings.TrimSpace(r.FormValue("comment_sort"))
p.CommentExtractorArgs = strings.TrimSpace(r.FormValue("comment_extractor_args"))
p.MaxComments = 0
if raw := strings.TrimSpace(r.FormValue("max_comments")); raw != "" {
maxComments, err := strconv.Atoi(raw)
if err != nil || maxComments < 0 {
return fmt.Errorf("Max comments must be a non-negative number.")
}
p.MaxComments = maxComments
}
// Comments are stored in the info JSON sidecar. Keep the dependent options
// consistent even when a client submits the form without JavaScript.
if !p.WriteInfoJSON || !p.WriteComments {
p.WriteComments = false
p.CommentSort = ""
p.MaxComments = 0
p.CommentExtractorArgs = ""
}
return nil
}
func (h *Handler) UpdatePreset(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
preset, err := h.presetSvc.GetByID(id)
if err != nil {
http.Error(w, "Not found", http.StatusNotFound)
return
}
if err := applyPresetForm(preset, r); err != nil {
redirectWithError(w, r, "/settings", err.Error(), nil)
return
}
if err := h.presetSvc.Update(preset); err != nil {
redirectWithError(w, r, "/settings", "Couldn't update this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset updated.")
}
func (h *Handler) DeletePreset(w http.ResponseWriter, r *http.Request) {
id, ok := parseID(w, r)
if !ok {
return
}
if err := h.presetSvc.Delete(id); err != nil {
redirectWithError(w, r, "/settings", "Couldn't delete this preset.", err)
return
}
redirectWithSuccess(w, r, "/settings", "Preset deleted.")
}
func (h *Handler) UpdateSettings(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
var firstErr error
record := func(err error) {
if err != nil && firstErr == nil {
firstErr = err
}
}
if interval := r.FormValue("refresh_interval"); interval != "" {
record(h.settingsSvc.SetRefreshInterval(interval))
}
record(h.settingsSvc.SetAutoRefreshLibrary(r.FormValue("auto_refresh_library") == "1"))
record(h.settingsSvc.SetAutoRefreshDownloads(r.FormValue("auto_refresh_downloads") == "1"))
record(h.settingsSvc.SetCookies(r.FormValue("cookies")))
if firstErr != nil {
log.Printf("UpdateSettings: %v", firstErr)
flashError(w, "Some settings couldn't be saved: "+firstErr.Error())
} else {
flashSuccess(w, "Settings saved.")
}
http.Redirect(w, r, "/settings", http.StatusSeeOther)
}
func (h *Handler) Theme(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
theme := r.FormValue("theme")
if theme == "" {
theme = "auto"
}
setCookie(w, "theme", theme)
referer := r.Header.Get("Referer")
if referer == "" {
referer = "/"
}
http.Redirect(w, r, referer, http.StatusSeeOther)
}
Minternal/handler/subscription.go
@@ -2,6 +2,7 @@ package handler
import (
"database/sql"
"log"
"net/http"
"strconv"
"strings"
@@ -175,12 +176,33 @@ func (h *Handler) RunSubscription(w http.ResponseWriter, r *http.Request) {
http.Error(w, "Not found", http.StatusNotFound)
return
}
// Refuse a second run while one is still in flight, the same guard the
// scheduler applies. Two runs of one subscription share an output directory,
// so in overwrite mode they race: one deletes the item the other just wrote.
active, err := h.downloadSvc.HasActiveForSubscription(id)
if err != nil {
redirectWithError(w, r, "/subscriptions", "Couldn't start this subscription run.", err)
return
}
if active {
redirectWithError(w, r, "/subscriptions", "This subscription already has a run in progress.", nil)
return
}
download, err := h.downloadSvc.CreateForSubscription(sub)
if err != nil {
redirectWithError(w, r, "/subscriptions", "Couldn't start this subscription run.", err)
return
}
h.workerPool.Submit(download)
// Record the manual run so the subscriptions page shows it. The schedule's
// next run time is deliberately left alone.
if err := h.subscriptionSvc.MarkManualRun(id, time.Now(), "queued"); err != nil {
log.Printf("subscription %d: failed to record manual run: %v", id, err)
}
redirectWithSuccess(w, r, "/queue", "Subscription run queued.")
}
Ainternal/repository/download_queue_test.go
@@ -0,0 +1,244 @@
package repository
import (
"context"
"database/sql"
"testing"
"vidarchive/internal/models"
)
func queueDownload(t *testing.T, repo *DownloadRepository, url, status string, subID *int64) *models.Download {
t.Helper()
d := &models.Download{URL: url, Status: status}
if subID != nil {
d.SubscriptionID = sql.NullInt64{Int64: *subID, Valid: true}
}
if err := repo.Create(d); err != nil {
t.Fatalf("create %s: %v", url, err)
}
return d
}
// The queue checker asks for the oldest queued rows up to the worker buffer
// size. Anything else, in any other status, must stay out of the result.
func TestGetQueuedOldestFirstWithinLimit(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
first := queueDownload(t, repo, "first", "queued", nil)
second := queueDownload(t, repo, "second", "queued", nil)
queueDownload(t, repo, "third", "queued", nil)
queueDownload(t, repo, "running", "downloading", nil)
queueDownload(t, repo, "done", "completed", nil)
got, err := repo.GetQueued(2)
if err != nil {
t.Fatalf("get queued: %v", err)
}
if len(got) != 2 {
t.Fatalf("got %d rows, want 2 (the limit)", len(got))
}
if got[0].ID != first.ID || got[1].ID != second.ID {
t.Errorf("got ids %d,%d, want the two oldest %d,%d", got[0].ID, got[1].ID, first.ID, second.ID)
}
}
// The queue page filters by status; "" and "all" mean no filter.
func TestGetAllStatusFilter(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
queueDownload(t, repo, "a", "queued", nil)
queueDownload(t, repo, "b", "error", nil)
queueDownload(t, repo, "c", "completed", nil)
for _, tc := range []struct {
status string
want int
}{
{"", 3},
{"all", 3},
{"error", 1},
{"cancelled", 0},
} {
got, err := repo.GetAll(tc.status, "date")
if err != nil {
t.Fatalf("get all %q: %v", tc.status, err)
}
if len(got) != tc.want {
t.Errorf("status %q returned %d rows, want %d", tc.status, len(got), tc.want)
}
}
// Sorting by status must not drop rows.
got, err := repo.GetAll("", "status")
if err != nil {
t.Fatalf("get all sorted by status: %v", err)
}
if len(got) != 3 {
t.Errorf("status sort returned %d rows, want 3", len(got))
}
}
// The health endpoint reads these counts, aggregated in SQL so it never loads
// the log column.
func TestCountByStatus(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
queueDownload(t, repo, "a", "queued", nil)
queueDownload(t, repo, "b", "queued", nil)
queueDownload(t, repo, "c", "error", nil)
counts, err := repo.CountByStatus(context.Background())
if err != nil {
t.Fatalf("count by status: %v", err)
}
if counts["queued"] != 2 || counts["error"] != 1 {
t.Errorf("counts = %v, want queued=2 error=1", counts)
}
if _, ok := counts["completed"]; ok {
t.Errorf("counts = %v, want no entry for an unused status", counts)
}
}
// Restart recovery: the stalled downloads are found by status and moved back to
// queued in one statement.
func TestIDsByStatusAndUpdateStatusWhere(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
stalled := queueDownload(t, repo, "stalled", "downloading", nil)
queueDownload(t, repo, "done", "completed", nil)
ids, err := repo.IDsByStatus("downloading")
if err != nil {
t.Fatalf("ids by status: %v", err)
}
if len(ids) != 1 || ids[0] != stalled.ID {
t.Fatalf("ids = %v, want [%d]", ids, stalled.ID)
}
if err := repo.UpdateStatusWhere("downloading", "queued"); err != nil {
t.Fatalf("update status: %v", err)
}
if got, _ := repo.GetByID(stalled.ID); got.Status != "queued" {
t.Errorf("status = %q, want queued", got.Status)
}
if left, _ := repo.IDsByStatus("downloading"); len(left) != 0 {
t.Errorf("%d rows still downloading", len(left))
}
}
// Logs are flushed in slices as yt-dlp produces them, so appends must
// accumulate rather than replace — including the first one, onto a NULL column.
func TestAppendLogsAccumulates(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
d := queueDownload(t, repo, "a", "downloading", nil)
for _, chunk := range []string{"first\n", "second\n"} {
if err := repo.AppendLogs(d.ID, chunk); err != nil {
t.Fatalf("append %q: %v", chunk, err)
}
}
got, err := repo.GetByID(d.ID)
if err != nil {
t.Fatal(err)
}
if got.Logs.String != "first\nsecond\n" {
t.Errorf("logs = %q, want both chunks in order", got.Logs.String)
}
}
func TestMarkErrorAndCompleted(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
failed := queueDownload(t, repo, "bad", "downloading", nil)
if err := repo.MarkError(failed.ID, "yt-dlp exploded"); err != nil {
t.Fatalf("mark error: %v", err)
}
got, _ := repo.GetByID(failed.ID)
if got.Status != "error" || got.ErrorMessage.String != "yt-dlp exploded" {
t.Errorf("got status=%q message=%q", got.Status, got.ErrorMessage.String)
}
ok := queueDownload(t, repo, "good", "downloading", nil)
if err := repo.MarkCompleted(ok.ID, "cancelled"); err != nil {
t.Fatalf("mark completed: %v", err)
}
got, _ = repo.GetByID(ok.ID)
if got.Status != "cancelled" || !got.CompletedAt.Valid {
t.Errorf("got status=%q completed_at valid=%v", got.Status, got.CompletedAt.Valid)
}
}
// The scheduler and the manual run button both use this to avoid stacking a
// second run on a subscription that is still working.
func TestHasActiveForSubscription(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
subRepo := NewSubscriptionRepository(db)
repo := NewDownloadRepository(db)
sub := &models.Subscription{
Name: "s", URL: "u", Enabled: true, RefreshMode: "overwrite",
ScheduleKind: "daily", CronExpr: "0 3 * * *", OutputDir: "dir",
}
if err := subRepo.Create(sub); err != nil {
t.Fatal(err)
}
if active, err := repo.HasActiveForSubscription(sub.ID); err != nil || active {
t.Fatalf("active = %v (err %v), want false with no downloads", active, err)
}
for _, status := range []string{"queued", "downloading"} {
d := queueDownload(t, repo, "u", status, &sub.ID)
active, err := repo.HasActiveForSubscription(sub.ID)
if err != nil {
t.Fatal(err)
}
if !active {
t.Errorf("status %q should count as active", status)
}
if err := repo.Delete(d.ID); err != nil {
t.Fatal(err)
}
}
// A finished run no longer blocks the next one.
queueDownload(t, repo, "u", "completed", &sub.ID)
if active, _ := repo.HasActiveForSubscription(sub.ID); active {
t.Error("a completed run must not count as active")
}
}
func TestDeleteAllClearsQueue(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewDownloadRepository(db)
queueDownload(t, repo, "a", "queued", nil)
queueDownload(t, repo, "b", "completed", nil)
if err := repo.DeleteAll(); err != nil {
t.Fatalf("delete all: %v", err)
}
got, err := repo.GetAll("", "date")
if err != nil {
t.Fatal(err)
}
if len(got) != 0 {
t.Errorf("%d rows left after DeleteAll", len(got))
}
}
Ainternal/repository/settings_test.go
@@ -0,0 +1,100 @@
package repository
import "testing"
// A fresh database is seeded by migration, so GetAll must report those values
// rather than the struct's zero value.
func TestSettingsRepositoryDefaults(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
settings, err := NewSettingsRepository(db).GetAll()
if err != nil {
t.Fatalf("get all: %v", err)
}
if settings.RefreshInterval != 5 {
t.Errorf("RefreshInterval = %d, want 5", settings.RefreshInterval)
}
if settings.AutoRefreshLibrary {
t.Error("AutoRefreshLibrary should default to off")
}
if !settings.AutoRefreshDownloads {
t.Error("AutoRefreshDownloads should default to on")
}
if settings.Cookies != "" {
t.Errorf("Cookies = %q, want empty", settings.Cookies)
}
}
func TestSettingsRepositoryRoundTrip(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewSettingsRepository(db)
for key, value := range map[string]string{
"refresh_interval": "12",
"auto_refresh_library": "1",
"auto_refresh_downloads": "0",
"cookies": "# Netscape HTTP Cookie File\n",
} {
// Set twice: the second call must update in place rather than conflict on
// the primary key.
if err := repo.Set(key, "placeholder"); err != nil {
t.Fatalf("set %s: %v", key, err)
}
if err := repo.Set(key, value); err != nil {
t.Fatalf("update %s: %v", key, err)
}
}
settings, err := repo.GetAll()
if err != nil {
t.Fatalf("get all: %v", err)
}
if settings.RefreshInterval != 12 {
t.Errorf("RefreshInterval = %d, want 12", settings.RefreshInterval)
}
if !settings.AutoRefreshLibrary {
t.Error("AutoRefreshLibrary = false, want true")
}
if settings.AutoRefreshDownloads {
t.Error("AutoRefreshDownloads = true, want false")
}
if got, err := repo.Get("cookies"); err != nil || got != "# Netscape HTTP Cookie File\n" {
t.Errorf("Get(cookies) = %q, %v", got, err)
}
}
// A non-numeric interval must not overwrite the default with zero, which would
// make the page refresh in a loop.
func TestSettingsRepositoryIgnoresUnparseableInterval(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewSettingsRepository(db)
if err := repo.Set("refresh_interval", "not a number"); err != nil {
t.Fatal(err)
}
settings, err := repo.GetAll()
if err != nil {
t.Fatal(err)
}
if settings.RefreshInterval != 5 {
t.Errorf("RefreshInterval = %d, want the default 5", settings.RefreshInterval)
}
}
// A key that was never written reads as empty, not as an error: the cookies
// lookup runs on every download.
func TestSettingsRepositoryGetMissingKey(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
got, err := NewSettingsRepository(db).Get("nothing_here")
if err != nil {
t.Errorf("Get on a missing key returned %v, want no error", err)
}
if got != "" {
t.Errorf("Get = %q, want empty", got)
}
}
Minternal/repository/subscription.go
@@ -90,7 +90,8 @@ func (r *SubscriptionRepository) SetEnabled(id int64, enabled bool) error {
return err
}
// MarkRun records a run's outcome and the computed next run time.
// MarkRun records that a scheduled run was started, together with the next run
// time the scheduler computed for it.
func (r *SubscriptionRepository) MarkRun(id int64, lastRunAt, nextRunAt time.Time, status string) error {
_, err := r.db.Exec(
`UPDATE subscriptions SET last_run_at = ?, next_run_at = ?, last_status = ? WHERE id = ?`,
@@ -99,6 +100,25 @@ func (r *SubscriptionRepository) MarkRun(id int64, lastRunAt, nextRunAt time.Tim
return err
}
// MarkManualRun records a run started by hand. Unlike MarkRun it leaves
// next_run_at alone: running a subscription now must not move its schedule.
func (r *SubscriptionRepository) MarkManualRun(id int64, lastRunAt time.Time, status string) error {
_, err := r.db.Exec(
`UPDATE subscriptions SET last_run_at = ?, last_status = ? WHERE id = ?`,
lastRunAt, status, id,
)
return err
}
// SetLastStatus records the outcome of a run that has already started, without
// touching the run times. The download worker calls this as the run progresses,
// so the subscription reflects what actually happened rather than staying on the
// status it had when it was queued.
func (r *SubscriptionRepository) SetLastStatus(id int64, status string) error {
_, err := r.db.Exec(`UPDATE subscriptions SET last_status = ? WHERE id = ?`, status, id)
return err
}
func (r *SubscriptionRepository) Delete(id int64) error {
_, err := r.db.Exec(`DELETE FROM subscriptions WHERE id = ?`, id)
return err
Minternal/repository/subscription_test.go
@@ -125,3 +125,76 @@ func TestSubscriptionRepositoryGetDueAndMarkRun(t *testing.T) {
func nullTime(t time.Time) sql.NullTime {
return sql.NullTime{Time: t, Valid: true}
}
// The download worker reports a run's outcome through SetLastStatus. It must
// change only the status: the run times belong to whoever started the run.
func TestSetLastStatusKeepsRunTimes(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewSubscriptionRepository(db)
now := time.Now().Truncate(time.Second)
next := now.Add(24 * time.Hour)
sub := &models.Subscription{
Name: "s", URL: "u", Enabled: true, RefreshMode: "overwrite",
ScheduleKind: "daily", CronExpr: "0 3 * * *", OutputDir: "dir",
}
if err := repo.Create(sub); err != nil {
t.Fatal(err)
}
if err := repo.MarkRun(sub.ID, now, next, "queued"); err != nil {
t.Fatal(err)
}
if err := repo.SetLastStatus(sub.ID, "completed"); err != nil {
t.Fatalf("set last status: %v", err)
}
got, err := repo.GetByID(sub.ID)
if err != nil {
t.Fatal(err)
}
if got.LastStatus.String != "completed" {
t.Errorf("LastStatus = %q, want completed", got.LastStatus.String)
}
if !got.NextRunAt.Valid || !got.NextRunAt.Time.Equal(next) {
t.Errorf("NextRunAt = %v, want it unchanged at %v", got.NextRunAt.Time, next)
}
if !got.LastRunAt.Valid || !got.LastRunAt.Time.Equal(now) {
t.Errorf("LastRunAt = %v, want it unchanged at %v", got.LastRunAt.Time, now)
}
}
// Running a subscription by hand records the run but must not move its
// schedule: the next scheduled run stays where the cron expression put it.
func TestMarkManualRunKeepsNextRun(t *testing.T) {
db := setupTestDB(t)
defer db.Close()
repo := NewSubscriptionRepository(db)
scheduled := time.Now().Add(6 * time.Hour).Truncate(time.Second)
sub := &models.Subscription{
Name: "s", URL: "u", Enabled: true, RefreshMode: "overwrite",
ScheduleKind: "daily", CronExpr: "0 3 * * *", OutputDir: "dir",
NextRunAt: nullTime(scheduled),
}
if err := repo.Create(sub); err != nil {
t.Fatal(err)
}
manualAt := time.Now().Truncate(time.Second)
if err := repo.MarkManualRun(sub.ID, manualAt, "queued"); err != nil {
t.Fatalf("mark manual run: %v", err)
}
got, err := repo.GetByID(sub.ID)
if err != nil {
t.Fatal(err)
}
if got.LastStatus.String != "queued" || !got.LastRunAt.Time.Equal(manualAt) {
t.Errorf("run not recorded: status=%q last=%v", got.LastStatus.String, got.LastRunAt.Time)
}
if !got.NextRunAt.Time.Equal(scheduled) {
t.Errorf("NextRunAt = %v, want it unchanged at %v", got.NextRunAt.Time, scheduled)
}
}
Minternal/server/server.go
@@ -80,8 +80,6 @@ func (s *Server) setupRoutes() {
s.router.Post("/settings", s.handler.UpdateSettings)
s.router.Post("/theme", s.handler.Theme)
s.router.Get("/api/presets/{id}/flags", s.handler.GetPresetFlags)
}
func (s *Server) securityHeaders(next http.Handler) http.Handler {
Ainternal/server/subscription_run_test.go
@@ -0,0 +1,114 @@
package server
import (
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
)
// createSubscription posts the create form and returns the id of the new row.
// The test server's worker pool is never started, so a queued run stays queued —
// which is exactly the state the duplicate-run guard looks for.
func createSubscription(t *testing.T, router http.Handler) string {
t.Helper()
w := postForm(router, "/subscriptions", url.Values{
"name": {"My Sub"},
"url": {"https://example.com/playlist"},
"output_dir": {"subs/mine"},
"refresh_mode": {"overwrite"},
"schedule_kind": {"daily"},
})
if w.Code != http.StatusSeeOther {
t.Fatalf("create subscription: expected 303, got %d", w.Code)
}
// The delete form on the listing carries the id.
req := httptest.NewRequest("GET", "/subscriptions", nil)
lw := httptest.NewRecorder()
router.ServeHTTP(lw, req)
body := lw.Body.String()
const marker = `action="/subscriptions/`
i := strings.Index(body, marker)
if i < 0 {
t.Fatal("no subscription id found on the listing")
}
rest := body[i+len(marker):]
id := rest[:strings.IndexAny(rest, `/"`)]
if id == "" {
t.Fatal("empty subscription id")
}
return id
}
// Running a subscription by hand queues one run and records it.
func TestRunSubscriptionQueuesOnce(t *testing.T) {
srv, _, cleanup := setupTestServer(t)
defer cleanup()
router := srv.Router()
id := createSubscription(t, router)
w := postForm(router, "/subscriptions/"+id+"/run", nil)
if w.Code != http.StatusSeeOther {
t.Fatalf("run: expected 303, got %d", w.Code)
}
if hasErrorFlash(w) {
t.Error("first run reported an error")
}
// The queue now holds the run, and the subscription reports it.
req := httptest.NewRequest("GET", "/queue", nil)
qw := httptest.NewRecorder()
router.ServeHTTP(qw, req)
if !strings.Contains(qw.Body.String(), "https://example.com/playlist") {
t.Error("queue page does not show the subscription run")
}
req = httptest.NewRequest("GET", "/subscriptions", nil)
sw := httptest.NewRecorder()
router.ServeHTTP(sw, req)
if !strings.Contains(sw.Body.String(), "queued") {
t.Error("subscriptions page does not show the run status")
}
}
// A second "Run now" while the first is still queued must be refused. Both runs
// share the subscription's output directory, so in overwrite mode they would
// race: one deletes the item the other just imported.
func TestRunSubscriptionRefusesConcurrentRun(t *testing.T) {
srv, _, cleanup := setupTestServer(t)
defer cleanup()
router := srv.Router()
id := createSubscription(t, router)
if w := postForm(router, "/subscriptions/"+id+"/run", nil); hasErrorFlash(w) {
t.Fatal("first run reported an error")
}
second := postForm(router, "/subscriptions/"+id+"/run", nil)
if second.Code != http.StatusSeeOther {
t.Fatalf("second run: expected 303, got %d", second.Code)
}
if !hasErrorFlash(second) {
t.Error("second run was accepted while the first was still queued")
}
// Still exactly one queued download for that URL.
req := httptest.NewRequest("GET", "/queue", nil)
qw := httptest.NewRecorder()
router.ServeHTTP(qw, req)
if n := strings.Count(qw.Body.String(), "https://example.com/playlist"); n != 1 {
t.Errorf("queue shows %d runs, want 1", n)
}
}
func TestRunUnknownSubscriptionIsNotFound(t *testing.T) {
srv, _, cleanup := setupTestServer(t)
defer cleanup()
w := postForm(srv.Router(), "/subscriptions/4242/run", nil)
if w.Code != http.StatusNotFound {
t.Errorf("expected 404 for an unknown subscription, got %d", w.Code)
}
}
Minternal/service/download.go
@@ -2,30 +2,23 @@ package service
import (
"bufio"
"bytes"
"context"
"database/sql"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"os"
"os/exec"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"syscall"
"time"
"github.com/gabriel-vasile/mimetype"
"vidarchive/internal/config"
"vidarchive/internal/models"
"vidarchive/internal/repository"
"vidarchive/internal/util"
)
type DownloadService struct {
@@ -223,21 +216,6 @@ func (s *DownloadService) tempDirsFor(id int64) []string {
return []string{s.tempDirFor(id), s.tempNewDirFor(id)}
}
func (s *DownloadService) ListFormats(url string) ([]*models.FormatInfo, error) {
// Use machine-readable JSON (-J) rather than scraping the human "-F" table,
// whose columns/separators shift between yt-dlp versions. stderr is captured
// separately so warnings can't corrupt the JSON on stdout.
cmd := exec.Command(s.cfg.YTDLPPath, "-J", "--no-warnings", url)
var stderr bytes.Buffer
cmd.Stderr = &stderr
output, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("yt-dlp -J failed: %w\n%s", err, stderr.String())
}
return parseFormatJSON(output)
}
// ExecuteDownload runs the download for d. The bool reports whether this call
// actually processed it: false means another worker already claimed it (Submit
// and the queue checker can both enqueue the same row within the 2s poll window),
@@ -293,11 +271,15 @@ func (s *DownloadService) ExecuteDownload(parent context.Context, d *models.Down
isSubscription := d.SubscriptionID.Valid
for _, flags := range []string{d.CustomFlags, preset.CustomFlags} {
if err := checkReservedFlags(flags, isSubscription); err != nil {
s.finalizeError(d.ID, err)
s.finalizeError(d, err)
return false, err
}
}
// From here the run is really under way, so a subscription shows "downloading"
// instead of the "queued" the scheduler recorded.
s.recordSubscriptionStatus(d, "downloading")
tempDownloadDir := s.tempDirFor(d.ID)
if err := os.MkdirAll(tempDownloadDir, 0755); err != nil {
return false, fmt.Errorf("create temp download dir: %w", err)
@@ -345,10 +327,10 @@ func (s *DownloadService) ExecuteDownload(parent context.Context, d *models.Down
// lands just after yt-dlp exited 0 leaves runErr nil, and the item must not
// reach the library after the user removed it.
if ctx.Err() != nil {
return false, s.finalizeCancelled(parent, d.ID)
return false, s.finalizeCancelled(parent, d)
}
if runErr != nil {
s.finalizeError(d.ID, runErr)
s.finalizeError(d, runErr)
return false, runErr
}
@@ -386,16 +368,17 @@ func (s *DownloadService) ExecuteDownload(parent context.Context, d *models.Down
// Same cancellation check as above: a stop during post-processing must not
// be recorded as "completed".
if ctx.Err() != nil {
return false, s.finalizeCancelled(parent, d.ID)
return false, s.finalizeCancelled(parent, d)
}
if postErr != nil {
s.finalizeError(d.ID, postErr)
s.finalizeError(d, postErr)
return false, postErr
}
if err := s.repo.MarkCompleted(d.ID, "completed"); err != nil {
return false, err
}
s.recordSubscriptionStatus(d, "completed")
return true, nil
}
@@ -496,10 +479,25 @@ func (s *DownloadService) writeCookiesFile(cookies string) (string, error) {
return tmpFile.Name(), nil
}
func (s *DownloadService) finalizeError(id int64, err error) {
s.flushLogs(id)
if markErr := s.repo.MarkError(id, err.Error()); markErr != nil {
log.Printf("download %d: failed to record error: %v", id, markErr)
func (s *DownloadService) finalizeError(d *models.Download, err error) {
s.flushLogs(d.ID)
if markErr := s.repo.MarkError(d.ID, err.Error()); markErr != nil {
log.Printf("download %d: failed to record error: %v", d.ID, markErr)
}
s.recordSubscriptionStatus(d, "error")
}
// recordSubscriptionStatus mirrors a subscription download's state onto the
// subscription row. Without it the row keeps the status it had when the
// scheduler queued it, so the subscriptions page reports "queued" long after the
// run finished — or failed.
func (s *DownloadService) recordSubscriptionStatus(d *models.Download, status string) {
if !d.SubscriptionID.Valid || s.subscriptionSvc == nil {
return
}
if err := s.subscriptionSvc.SetLastStatus(d.SubscriptionID.Int64, status); err != nil {
log.Printf("download %d: failed to record subscription %d status %q: %v",
d.ID, d.SubscriptionID.Int64, status, err)
}
}
@@ -515,16 +513,19 @@ var ErrCancelled = errors.New("download cancelled")
// stopping the server resumes the download instead of losing it. Only a
// user-initiated cancel is terminal. The row may already be deleted in that
// case — cancellation usually arrives via Delete — so a missing row is fine.
func (s *DownloadService) finalizeCancelled(parent context.Context, id int64) error {
s.flushLogs(id)
func (s *DownloadService) finalizeCancelled(parent context.Context, d *models.Download) error {
s.flushLogs(d.ID)
// A shutdown leaves the subscription status alone too: the run resumes on the
// next start, so it is still in progress rather than cancelled.
if parent.Err() != nil {
return ErrCancelled
}
if err := s.repo.MarkCompleted(id, "cancelled"); err != nil {
log.Printf("download %d: failed to record cancellation: %v", id, err)
if err := s.repo.MarkCompleted(d.ID, "cancelled"); err != nil {
log.Printf("download %d: failed to record cancellation: %v", d.ID, err)
}
s.recordSubscriptionStatus(d, "cancelled")
return ErrCancelled
}
@@ -539,784 +540,6 @@ func (s *DownloadService) flushLogs(id int64) {
}
}
// resolveBaseLibraryDir returns the absolute library directory a download writes
// into, applying the optional per-download OutputDir while rejecting any path
// that escapes the library root.
func (s *DownloadService) resolveBaseLibraryDir(d *models.Download) (string, error) {
if !d.OutputDir.Valid || d.OutputDir.String == "" {
return s.cfg.LibraryDir, nil
}
// Reuse the library service's guard so both entry points enforce the boundary
// the same way — it resolves symlinks, which a plain prefix check does not.
dir, err := s.librarySvc.ResolveWithinLibrary(d.OutputDir.String)
if err != nil {
return "", fmt.Errorf("invalid output directory: %w", err)
}
return dir, nil
}
// importDownloadedItems moves each downloaded item from the temp dir into the
// library and returns the number of items successfully imported. Per-item
// failures are logged and skipped (a playlist with a few bad entries still
// imports the rest); a non-nil error means the import couldn't even start.
//
// A cancel stops the import between items and returns ctx.Err() with the count
// imported so far. Importing a long playlist takes real time (a move plus an
// ffprobe per file), so a deleted download must not keep filling the library.
func (s *DownloadService) importDownloadedItems(ctx context.Context, d *models.Download, tempDownloadDir, mode, ytdlpFlags string) (int, error) {
entries, err := os.ReadDir(tempDownloadDir)
if err != nil {
return 0, err
}
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
return 0, err
}
if err := os.MkdirAll(baseLibraryDir, 0755); err != nil {
return 0, err
}
var itemDirs []string
for _, entry := range entries {
if !entry.IsDir() {
continue
}
name := entry.Name()
if strings.HasPrefix(name, "item-") {
itemDirs = append(itemDirs, filepath.Join(tempDownloadDir, name))
}
}
sort.Strings(itemDirs)
imported := 0
for _, itemDir := range itemDirs {
if err := ctx.Err(); err != nil {
return imported, err
}
if err := s.importItemDir(ctx, d.URL, itemDir, baseLibraryDir, mode, ytdlpFlags); err != nil {
// A cancelled item isn't a bad item: stop instead of logging a warning
// for it and every one that follows.
if ctx.Err() != nil {
return imported, ctx.Err()
}
log.Printf("warning: failed to import item %s: %v", itemDir, err)
continue
}
imported++
}
// The temp dir (and any leftovers from failed imports) is removed by the
// caller's deferred cleanup, so partial state never leaks even on a crash.
return imported, nil
}
func (s *DownloadService) importItemDir(ctx context.Context, url, itemDir, baseLibraryDir, mode, ytdlpFlags string) error {
entries, err := os.ReadDir(itemDir)
if err != nil {
return err
}
var mediaFiles []os.DirEntry
var infoJSONPath string
var subtitleFiles []string
for _, entry := range entries {
if entry.IsDir() {
continue
}
name := entry.Name()
path := filepath.Join(itemDir, name)
ext := strings.ToLower(filepath.Ext(name))
if isInfoJSON(name) {
infoJSONPath = path
continue
}
if ext == ".vtt" || ext == ".srt" || ext == ".ass" || ext == ".ssa" {
subtitleFiles = append(subtitleFiles, path)
continue
}
mtype, err := mimetype.DetectFile(path)
if err == nil && mtype != nil && (strings.HasPrefix(mtype.String(), "audio/") || strings.HasPrefix(mtype.String(), "video/")) {
mediaFiles = append(mediaFiles, entry)
}
}
if len(mediaFiles) == 0 {
return fmt.Errorf("no media files found in %s", itemDir)
}
info := readInfoJSON(infoJSONPath)
name := s.deriveItemName(itemDir, info, mediaFiles)
videoID := info.ID
// Last point at which nothing has been written to the library yet: give up
// here on a cancel rather than part-way through, which would leave a folder
// with some of its files and no marker — or, in overwrite mode, delete the
// existing item and not replace it.
if err := ctx.Err(); err != nil {
return err
}
// Overwrite mode: replace the existing copy of this video in place rather than
// creating a duplicate folder. Removing the old dir lets uniqueDir reuse its
// name (or land on the new title if it changed upstream).
if mode == "overwrite" && videoID != "" {
if existing, ok := s.librarySvc.FindByVideoID(baseLibraryDir, videoID); ok {
if rel, err := filepath.Rel(s.cfg.LibraryDir, existing); err == nil {
s.librarySvc.evictCachedScan(filepath.ToSlash(rel))
}
os.RemoveAll(existing)
}
}
targetDir := s.uniqueDir(baseLibraryDir, name)
if err := os.MkdirAll(targetDir, 0755); err != nil {
return err
}
if infoJSONPath != "" {
if err := moveFile(infoJSONPath, filepath.Join(targetDir, "info.json")); err != nil {
return err
}
}
for _, entry := range mediaFiles {
if err := moveFile(filepath.Join(itemDir, entry.Name()), filepath.Join(targetDir, entry.Name())); err != nil {
return err
}
}
// Probe each media file's duration once, here in the worker (off the request
// path), and cache it in the marker so the library never has to probe while
// serving pages. Files we can't probe simply get no duration.
fileDurations := make(map[string]int)
for _, entry := range mediaFiles {
if d, ok := probeDuration(ctx, s.cfg.FFprobePath, filepath.Join(targetDir, entry.Name())); ok {
fileDurations[entry.Name()] = d
}
}
if len(subtitleFiles) > 0 {
subtitlesDir := filepath.Join(targetDir, subtitlesDirName)
if err := os.MkdirAll(subtitlesDir, 0755); err != nil {
return err
}
for _, sf := range subtitleFiles {
if err := moveFile(sf, filepath.Join(subtitlesDir, filepath.Base(sf))); err != nil {
return err
}
}
}
metadata := models.ItemMetadata{
Name: name,
SourceURL: url,
VideoID: videoID,
YtdlpFlags: ytdlpFlags,
FileDurations: fileDurations,
}
return s.librarySvc.writeMetadata(targetDir, metadata)
}
// infoJSON is the subset of yt-dlp's info.json VidArchive reads. ID is the
// stable item identity; within a single subscription's own directory it is
// enough to match items, so the extractor is not needed.
type infoJSON struct {
ID string `json:"id"`
Title string `json:"title"`
Description string `json:"description"`
WebpageURL string `json:"webpage_url"`
}
// readInfoJSON parses an info.json. A missing, unreadable or malformed file
// yields a zero-value struct: every caller treats absent fields as "unknown"
// and falls back, so there is nothing to distinguish.
func readInfoJSON(infoJSONPath string) infoJSON {
var info infoJSON
if infoJSONPath == "" {
return info
}
data, err := os.ReadFile(infoJSONPath)
if err != nil {
return info
}
if err := json.Unmarshal(data, &info); err != nil {
log.Printf("ignoring malformed %s: %v", infoJSONPath, err)
return infoJSON{}
}
return info
}
// refreshAndAddNew handles a metadata-mode run. The main pass used
// --skip-download, so tempDownloadDir holds only info.json files. Existing
// library items have their markers refreshed in place; entries with no existing
// match are genuinely new and are downloaded as full items in a second pass.
func (s *DownloadService) refreshAndAddNew(ctx context.Context, d *models.Download, preset *models.Preset, tempDownloadDir, ytdlpFlags string) error {
// The library dir is not created here: a refresh that matches everything
// writes nothing, and importDownloadedItems creates it when a second pass
// actually has an item to add.
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
return err
}
entries, err := os.ReadDir(tempDownloadDir)
if err != nil {
return err
}
var newURLs []string
for _, entry := range entries {
if err := ctx.Err(); err != nil {
return err
}
if !entry.IsDir() || !strings.HasPrefix(entry.Name(), "item-") {
continue
}
itemDir := filepath.Join(tempDownloadDir, entry.Name())
infoJSONPath := findInfoJSON(itemDir)
if infoJSONPath == "" {
continue
}
info := readInfoJSON(infoJSONPath)
if info.ID == "" {
continue
}
if existing, ok := s.librarySvc.FindByVideoID(baseLibraryDir, info.ID); ok {
if err := s.applyMetadata(existing, info, infoJSONPath); err != nil {
log.Printf("warning: failed to refresh metadata for %s: %v", itemDir, err)
}
continue
}
if info.WebpageURL != "" {
newURLs = append(newURLs, info.WebpageURL)
}
}
if len(newURLs) == 0 {
return nil
}
return s.downloadFresh(ctx, d, preset, newURLs, ytdlpFlags)
}
// downloadFresh fetches the given item URLs as full downloads (media + info.json)
// and imports them into the download's library directory. Metadata mode uses this
// to add entries that don't exist in the library yet.
func (s *DownloadService) downloadFresh(ctx context.Context, d *models.Download, preset *models.Preset, urls []string, ytdlpFlags string) error {
tempDir := s.tempNewDirFor(d.ID)
if err := os.MkdirAll(tempDir, 0755); err != nil {
return err
}
defer os.RemoveAll(tempDir)
args := s.presetSvc.BuildArgs(preset, d.FormatOverride, d.CustomFlags)
args, cleanup := s.appendCookies(args)
defer cleanup()
args = append(args, "--write-info-json")
args = append(args, "-P", tempDir)
args = append(args, "-o", "item-%(autonumber)05d/%(title)s.%(ext)s")
args = append(args, urls...)
runErr := s.runYTDLP(ctx, d, args)
// A cancelled second pass has nothing worth importing.
if ctx.Err() != nil {
return runErr
}
// Import whatever succeeded even if some entries errored.
if _, err := s.importDownloadedItems(ctx, d, tempDir, "", ytdlpFlags); err != nil {
log.Printf("warning: failed to import new metadata-mode items: %v", err)
}
return runErr
}
// mergeInfoJSON keeps fields from the existing sidecar that are absent from a
// metadata-only refresh. In particular, comments and heatmap data are expensive
// to reacquire and must not disappear just because the refresh preset does not
// request them.
func mergeInfoJSON(oldData, newData []byte) ([]byte, error) {
var oldObject, newObject map[string]json.RawMessage
if err := json.Unmarshal(newData, &newObject); err != nil {
return nil, err
}
if err := json.Unmarshal(oldData, &oldObject); err != nil {
return newData, nil
}
if newObject == nil {
return newData, nil
}
merged := make(map[string]json.RawMessage, len(oldObject)+len(newObject))
for key, value := range oldObject {
merged[key] = value
}
for key, value := range newObject {
merged[key] = value
}
for _, key := range []string{"comments", "heatmap"} {
oldValue, hadOldValue := oldObject[key]
newValue, hasNewValue := newObject[key]
if hadOldValue && (!hasNewValue || isEmptyJSONArray(newValue)) {
merged[key] = oldValue
}
}
return json.Marshal(merged)
}
func isEmptyJSONArray(value json.RawMessage) bool {
var values []json.RawMessage
if err := json.Unmarshal(value, &values); err != nil {
return false
}
return len(values) == 0
}
// restoreFileAtomically puts data back at path without exposing a partial file.
// It is used to roll back the marker if installing the staged info sidecar fails.
func restoreFileAtomically(path string, data []byte) error {
tmp, err := os.CreateTemp(filepath.Dir(path), ".vidarchive-restore-*.tmp")
if err != nil {
return err
}
tmpPath := tmp.Name()
defer os.Remove(tmpPath)
if err := tmp.Chmod(0644); err != nil {
tmp.Close()
return err
}
if _, err := tmp.Write(data); err != nil {
tmp.Close()
return err
}
if err := tmp.Close(); err != nil {
return err
}
if err := os.Rename(tmpPath, path); err != nil {
return err
}
return nil
}
// applyMetadata refreshes an existing item's marker and info sidecar from a
// fresh info.json without touching its media.
func (s *DownloadService) applyMetadata(existing string, info infoJSON, sourceInfoJSON string) error {
meta, err := s.librarySvc.readMetadata(existing)
if err != nil {
return fmt.Errorf("read metadata for %s: %w", existing, err)
}
if info.Title != "" {
meta.Name = info.Title
}
if info.Description != "" {
meta.Description = info.Description
}
if meta.SourceURL == "" && info.WebpageURL != "" {
meta.SourceURL = info.WebpageURL
}
meta.VideoID = info.ID
markerPath := filepath.Join(existing, itemMarkerName)
oldMarker, err := os.ReadFile(markerPath)
if err != nil {
return fmt.Errorf("read existing marker: %w", err)
}
// Stage the sidecar before changing the marker. The marker is committed first;
// if installing the sidecar then fails, restore the old marker so an ordinary
// I/O error cannot leave the two metadata files out of sync.
stagedInfo := ""
defer func() {
if stagedInfo != "" {
_ = os.Remove(stagedInfo)
}
}()
if sourceInfoJSON != "" {
data, err := os.ReadFile(sourceInfoJSON)
if err != nil {
return fmt.Errorf("read refreshed info JSON: %w", err)
}
if oldInfoJSON := findInfoJSON(existing); oldInfoJSON != "" {
oldData, err := os.ReadFile(oldInfoJSON)
if err != nil {
return fmt.Errorf("read existing info JSON: %w", err)
}
data, err = mergeInfoJSON(oldData, data)
if err != nil {
return fmt.Errorf("merge refreshed info JSON: %w", err)
}
}
tmp, err := os.CreateTemp(existing, ".info-json-*.tmp")
if err != nil {
return fmt.Errorf("create refreshed info JSON: %w", err)
}
stagedInfo = tmp.Name()
if err := tmp.Chmod(0644); err != nil {
tmp.Close()
return fmt.Errorf("set refreshed info JSON permissions: %w", err)
}
if _, err := tmp.Write(data); err != nil {
tmp.Close()
return fmt.Errorf("write refreshed info JSON: %w", err)
}
if err := tmp.Close(); err != nil {
return fmt.Errorf("close refreshed info JSON: %w", err)
}
}
if err := s.librarySvc.writeMetadata(existing, meta); err != nil {
return err
}
if stagedInfo != "" {
if err := os.Rename(stagedInfo, filepath.Join(existing, "info.json")); err != nil {
if restoreErr := restoreFileAtomically(markerPath, oldMarker); restoreErr != nil {
return fmt.Errorf("install refreshed info JSON: %v; restore marker: %w", err, restoreErr)
}
return fmt.Errorf("install refreshed info JSON: %w", err)
}
stagedInfo = ""
}
if rel, err := filepath.Rel(s.cfg.LibraryDir, existing); err == nil {
s.librarySvc.evictCachedScan(filepath.ToSlash(rel))
}
return nil
}
// isInfoJSON reports whether a file name is yt-dlp's metadata sidecar. yt-dlp
// writes "<title>.info.json" next to the media, but a bare "info.json" is what
// an already-imported item holds.
func isInfoJSON(name string) bool {
return name == "info.json" || strings.HasSuffix(name, ".info.json")
}
// findInfoJSON returns the path to an info.json directly inside itemDir, or "".
func findInfoJSON(itemDir string) string {
entries, err := os.ReadDir(itemDir)
if err != nil {
return ""
}
for _, entry := range entries {
if entry.IsDir() {
continue
}
name := entry.Name()
if isInfoJSON(name) {
return filepath.Join(itemDir, name)
}
}
return ""
}
// pruneSubscription mirrors the source by deleting items in the subscription's
// directory that are no longer present upstream. It enumerates the current id
// set with a cheap flat-playlist listing; it never prunes when that enumeration
// fails or returns nothing, so a dead URL or network error can't wipe the dir.
func (s *DownloadService) pruneSubscription(ctx context.Context, d *models.Download, sub *models.Subscription) {
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
log.Printf("subscription %d prune skipped: %v", sub.ID, err)
return
}
keep, err := s.enumeratePlaylistIDs(ctx, sub.URL)
if err != nil {
log.Printf("subscription %d prune skipped: enumeration failed: %v", sub.ID, err)
return
}
if len(keep) == 0 {
log.Printf("subscription %d prune skipped: source returned no entries", sub.ID)
return
}
removed, err := s.librarySvc.PruneToIDSet(baseLibraryDir, keep)
if err != nil {
log.Printf("subscription %d prune error: %v", sub.ID, err)
return
}
if removed > 0 {
log.Printf("subscription %d pruned %d item(s) removed upstream", sub.ID, removed)
}
}
// enumeratePlaylistIDs lists the current video-id set for a URL without
// downloading, using yt-dlp --flat-playlist. Cookies are applied so private
// playlists enumerate correctly. Ids alone are sufficient to match items within
// a subscription's own directory (see FindByVideoID / PruneToIDSet).
func (s *DownloadService) enumeratePlaylistIDs(ctx context.Context, url string) (map[string]bool, error) {
args := []string{"--flat-playlist", "--no-warnings", "--print", "%(id)s"}
args, cleanup := s.appendCookies(args)
defer cleanup()
args = append(args, url)
out, err := exec.CommandContext(ctx, s.cfg.YTDLPPath, args...).Output()
if err != nil {
return nil, err
}
keep := make(map[string]bool)
for _, line := range strings.Split(string(out), "\n") {
id := strings.TrimSpace(line)
// yt-dlp prints "NA" for a missing field; never treat that as a real id.
if id == "" || id == "NA" {
continue
}
keep[id] = true
}
return keep, nil
}
// deriveItemName names the imported item after its title, falling back to the
// largest media file's base name when there is no usable info.json.
func (s *DownloadService) deriveItemName(itemDir string, info infoJSON, mediaFiles []os.DirEntry) string {
if info.Title != "" {
return sanitizeDirName(info.Title)
}
var largest os.DirEntry
var maxSize int64
for _, f := range mediaFiles {
st, err := os.Stat(filepath.Join(itemDir, f.Name()))
if err == nil && (largest == nil || st.Size() > maxSize) {
largest, maxSize = f, st.Size()
}
}
if largest == nil {
largest = mediaFiles[0]
}
base := strings.TrimSuffix(largest.Name(), filepath.Ext(largest.Name()))
return sanitizeDirName(base)
}
func (s *DownloadService) uniqueDir(base, name string) string {
dir := filepath.Join(base, name)
if _, err := os.Stat(dir); os.IsNotExist(err) {
return dir
}
for i := 1; ; i++ {
candidate := fmt.Sprintf("%s-%d", dir, i)
if _, err := os.Stat(candidate); os.IsNotExist(err) {
return candidate
}
}
}
// moveFile moves src to dst, falling back to copy-and-delete when the two are on
// different filesystems. The temp and library directories are independently
// configurable, so they can legitimately live on separate mounts — where a plain
// rename fails with EXDEV.
func moveFile(src, dst string) error {
if err := os.Rename(src, dst); err == nil {
return nil
} else if !errors.Is(err, syscall.EXDEV) {
return err
}
in, err := os.Open(src)
if err != nil {
return err
}
defer in.Close()
info, err := in.Stat()
if err != nil {
return err
}
out, err := os.OpenFile(dst, os.O_WRONLY|os.O_CREATE|os.O_TRUNC, info.Mode())
if err != nil {
return err
}
if _, err := io.Copy(out, in); err != nil {
out.Close()
os.Remove(dst)
return err
}
// Close explicitly: a deferred close would hide a flush error on the copy.
if err := out.Close(); err != nil {
os.Remove(dst)
return err
}
return os.Remove(src)
}
func sanitizeDirName(name string) string {
name = strings.TrimSpace(name)
replacer := strings.NewReplacer(
"/", "-",
"\\", "-",
":", "-",
"*", "-",
"?", "-",
"\"", "-",
"<", "-",
">", "-",
"|", "-",
)
name = replacer.Replace(name)
name = strings.TrimSpace(name)
if name == "" {
name = "untitled"
}
return name
}
// probeDuration returns the duration of a media file in whole seconds. The bool
// is false when ffprobe is unavailable or the file has no usable duration.
func probeDuration(ctx context.Context, ffprobePath, path string) (int, bool) {
out, err := exec.CommandContext(ctx, ffprobePath, "-v", "error",
"-show_entries", "format=duration",
"-of", "default=nw=1:nk=1", path).Output()
if err != nil {
return 0, false
}
f, err := strconv.ParseFloat(strings.TrimSpace(string(out)), 64)
if err != nil || f <= 0 {
return 0, false
}
return int(f + 0.5), true
}
// ytFormat mirrors the subset of yt-dlp's per-format JSON (-J) we surface.
// Numeric fields are pointers so an absent value (null/omitted) is distinct
// from a real zero.
type ytFormat struct {
FormatID string `json:"format_id"`
Ext string `json:"ext"`
Resolution string `json:"resolution"`
Width *int `json:"width"`
Height *int `json:"height"`
FPS *float64 `json:"fps"`
VCodec string `json:"vcodec"`
ACodec string `json:"acodec"`
AudioChannels *int `json:"audio_channels"`
Filesize *int64 `json:"filesize"`
FilesizeApprox *int64 `json:"filesize_approx"`
FormatNote string `json:"format_note"`
}
// parseFormatJSON reads yt-dlp's single-JSON dump (-J) and returns the available
// formats. For a single video the formats live at the top level; for a playlist
// URL we fall back to the first entry's formats so the picker still shows
// something useful.
func parseFormatJSON(data []byte) ([]*models.FormatInfo, error) {
var top struct {
Formats []ytFormat `json:"formats"`
Entries []struct {
Formats []ytFormat `json:"formats"`
} `json:"entries"`
}
if err := json.Unmarshal(data, &top); err != nil {
return nil, fmt.Errorf("parse yt-dlp JSON: %w", err)
}
raw := top.Formats
if len(raw) == 0 && len(top.Entries) > 0 {
raw = top.Entries[0].Formats
}
formats := make([]*models.FormatInfo, 0, len(raw))
for _, f := range raw {
formats = append(formats, f.toFormatInfo())
}
return formats, nil
}
func (f ytFormat) toFormatInfo() *models.FormatInfo {
fi := &models.FormatInfo{
ID: f.FormatID,
Ext: f.Ext,
Note: f.FormatNote,
}
switch {
case f.Resolution != "":
fi.Resolution = f.Resolution
case f.Width != nil && f.Height != nil && *f.Width > 0 && *f.Height > 0:
fi.Resolution = fmt.Sprintf("%dx%d", *f.Width, *f.Height)
}
if f.FPS != nil && *f.FPS > 0 {
fi.FPS = strconv.FormatFloat(*f.FPS, 'f', -1, 64)
}
if f.AudioChannels != nil && *f.AudioChannels > 0 {
fi.Channels = strconv.Itoa(*f.AudioChannels)
}
// Prefer the video codec; fall back to the audio codec for audio-only formats.
if f.VCodec != "" && f.VCodec != "none" {
fi.Codec = f.VCodec
} else if f.ACodec != "" && f.ACodec != "none" {
fi.Codec = f.ACodec
}
if f.Filesize != nil && *f.Filesize > 0 {
fi.FileSize = util.FormatBytes(*f.Filesize)
} else if f.FilesizeApprox != nil && *f.FilesizeApprox > 0 {
fi.FileSize = "~" + util.FormatBytes(*f.FilesizeApprox)
}
return fi
}
func sqlNullInt64(v int64) sql.NullInt64 {
return sql.NullInt64{Int64: v, Valid: true}
}
// reservedFlags are yt-dlp options VidArchive always sets itself; user custom
// flags must not pass them (or a conflicting inverse). The value describes what
// the option controls, for the failure message.
var reservedFlags = map[string]string{
"-o": "the output template",
"--output": "the output template",
"-P": "the download path",
"--paths": "the download path",
"--cookies": "cookies (set these in Settings instead)",
"--no-cookies": "cookies (set these in Settings instead)",
"--newline": "progress output formatting (VidArchive sets this to stream logs)",
// These hand yt-dlp an arbitrary command or binary to run. VidArchive passes
// custom flags through verbatim, so allowing them would turn the preset form
// into remote command execution.
"--exec": "running external commands (not permitted)",
"--exec-before-download": "running external commands (not permitted)",
"--postprocessor-args": "post-processor arguments (not permitted)",
"--ppa": "post-processor arguments (not permitted)",
"--downloader": "selecting an external downloader (not permitted)",
"--external-downloader": "selecting an external downloader (not permitted)",
"--downloader-args": "external downloader arguments (not permitted)",
"--external-downloader-args": "external downloader arguments (not permitted)",
}
// reservedSubscriptionFlags are additionally reserved for subscription runs,
// where VidArchive drives info-json writing and the refresh mode.
var reservedSubscriptionFlags = map[string]string{
"--write-info-json": "info-json writing (needed to track item identity)",
"--no-write-info-json": "info-json writing (needed to track item identity)",
"--download-archive": "the download archive (managed by Skip mode)",
"--no-download-archive": "the download archive (managed by Skip mode)",
"--skip-download": "media downloading (managed by Metadata mode)",
"--no-skip-download": "media downloading (managed by Metadata mode)",
}
// checkReservedFlags rejects custom flags that clash with options VidArchive
// controls, naming the offender. It matches both "--flag" and "--flag=value".
func checkReservedFlags(customFlags string, isSubscription bool) error {
for _, tok := range strings.Fields(customFlags) {
// Both "--flag value" and "--flag=value" name the same option.
name, _, _ := strings.Cut(tok, "=")
desc, ok := reservedFlags[name]
if !ok && isSubscription {
desc, ok = reservedSubscriptionFlags[name]
}
if ok {
return fmt.Errorf("custom flag %q conflicts with VidArchive's handling of %s; remove it and try again", tok, desc)
}
}
return nil
}
Ainternal/service/engagement.go
@@ -0,0 +1,206 @@
package service
import (
"encoding/json"
"fmt"
"math"
"os"
"sort"
"strings"
"time"
"vidarchive/internal/models"
)
// GetEngagement reads optional comments and a playback heatmap from the item's
// yt-dlp info sidecar. These fields are intentionally not copied into the
// marker: the sidecar remains the source of truth and old items simply return
// empty data when the fields are absent.
func (s *LibraryService) GetEngagement(relPath string) ([]models.Comment, []models.HeatmapSegment, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, nil, err
}
_, infoPath, err := s.listItemFiles(item.DirPath)
if err != nil || infoPath == "" {
return nil, nil, err
}
data, err := os.ReadFile(infoPath)
if err != nil {
return nil, nil, err
}
var raw struct {
Comments []struct {
ID string `json:"id"`
Parent string `json:"parent"`
Author string `json:"author"`
Channel string `json:"channel"`
Text string `json:"text"`
Timestamp float64 `json:"timestamp"`
TimeText string `json:"time_text"`
LikeCount int `json:"like_count"`
AuthorIsUploader bool `json:"author_is_uploader"`
} `json:"comments"`
Heatmap []struct {
StartTime float64 `json:"start_time"`
EndTime float64 `json:"end_time"`
Value float64 `json:"value"`
} `json:"heatmap"`
}
if err := json.Unmarshal(data, &raw); err != nil {
return nil, nil, fmt.Errorf("parse engagement info JSON: %w", err)
}
comments := make([]models.Comment, 0, len(raw.Comments))
for _, c := range raw.Comments {
author := c.Author
if author == "" {
author = c.Channel
}
if c.Text == "" && author == "" {
continue
}
comment := models.Comment{
ID: c.ID,
Parent: c.Parent,
Author: author,
Text: c.Text,
Timestamp: int64(c.Timestamp),
TimeText: c.TimeText,
LikeCount: c.LikeCount,
AuthorIsUploader: c.AuthorIsUploader,
}
if comment.TimeText == "" && comment.Timestamp > 0 {
comment.TimeText = time.Unix(comment.Timestamp, 0).UTC().Format("2006-01-02 15:04")
}
comments = append(comments, comment)
}
comments = orderComments(comments)
maxValue := 0.0
for _, h := range raw.Heatmap {
if h.Value > maxValue && !math.IsNaN(h.Value) && !math.IsInf(h.Value, 0) {
maxValue = h.Value
}
}
endTime := 0.0
for _, h := range raw.Heatmap {
if h.EndTime > endTime {
endTime = h.EndTime
}
}
heatmap := make([]models.HeatmapSegment, 0, len(raw.Heatmap))
for _, h := range raw.Heatmap {
if h.EndTime <= h.StartTime || h.StartTime < 0 || h.Value < 0 ||
math.IsNaN(h.StartTime) || math.IsNaN(h.EndTime) || math.IsNaN(h.Value) ||
math.IsInf(h.StartTime, 0) || math.IsInf(h.EndTime, 0) || math.IsInf(h.Value, 0) {
continue
}
width := 0.0
if endTime > 0 {
width = (h.EndTime - h.StartTime) / endTime * 100
}
height := 8.0
if maxValue > 0 {
height += h.Value / maxValue * 92
}
heatmap = append(heatmap, models.HeatmapSegment{
StartTime: h.StartTime,
EndTime: h.EndTime,
Value: h.Value,
Width: width,
Height: height,
})
}
sort.SliceStable(heatmap, func(i, j int) bool {
return heatmap[i].StartTime < heatmap[j].StartTime
})
return comments, heatmap, nil
}
// orderComments groups replies beneath their parent while preserving the
// source order among siblings. Depth is capped for presentation so malformed
// or unusually deep reply chains cannot make the UI progressively narrower.
func orderComments(comments []models.Comment) []models.Comment {
if len(comments) < 2 {
return comments
}
byID := make(map[string]int, len(comments))
for i, comment := range comments {
if comment.ID != "" {
if _, exists := byID[comment.ID]; !exists {
byID[comment.ID] = i
}
}
}
children := make(map[int][]int)
var roots []int
for i, comment := range comments {
parent := strings.TrimSpace(comment.Parent)
if parent == "" || strings.EqualFold(parent, "root") {
roots = append(roots, i)
continue
}
parentIndex, ok := byID[parent]
if !ok || parentIndex == i {
roots = append(roots, i)
continue
}
children[parentIndex] = append(children[parentIndex], i)
}
ordered := make([]models.Comment, 0, len(comments))
visited := make([]bool, len(comments))
stack := make([]struct {
index int
depth int
}, 0, len(comments))
walk := func(index, depth int) {
stack = append(stack, struct {
index int
depth int
}{index, depth})
for len(stack) > 0 {
last := len(stack) - 1
entry := stack[last]
stack = stack[:last]
if visited[entry.index] {
continue
}
visited[entry.index] = true
comment := comments[entry.index]
if parentIndex, ok := byID[strings.TrimSpace(comment.Parent)]; ok && parentIndex != entry.index {
comment.ReplyTo = comments[parentIndex].Author
}
if entry.depth > 4 {
comment.Depth = 4
} else {
comment.Depth = entry.depth
}
ordered = append(ordered, comment)
childrenForComment := children[entry.index]
for i := len(childrenForComment) - 1; i >= 0; i-- {
stack = append(stack, struct {
index int
depth int
}{childrenForComment[i], entry.depth + 1})
}
}
}
for _, root := range roots {
walk(root, 0)
}
// Cycles or references to invalid parents are rendered as top-level comments
// rather than being dropped.
for i := range comments {
if !visited[i] {
walk(i, 0)
}
}
return ordered
}
Ainternal/service/engagement_test.go
@@ -0,0 +1,136 @@
package service
import (
"reflect"
"testing"
"vidarchive/internal/models"
)
func TestGetEngagement(t *testing.T) {
svc, dir := newLibrary(t)
info := `{"comments":[{"author":"Alice","text":"hello <world>","timestamp":1700000000,"like_count":3,"author_is_uploader":true},{"channel":"Bob","text":"second","time_text":"yesterday"}],"heatmap":[{"start_time":0,"end_time":10,"value":0.5},{"start_time":10,"end_time":20,"value":1.0}]}`
writeItem(t, dir, "engagement", `name = "Engagement"`, map[string]string{
"video.mp4": "dummy",
"info.json": info,
})
comments, heatmap, err := svc.GetEngagement("engagement")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(comments) != 2 || comments[0].Author != "Alice" || comments[0].LikeCount != 3 || !comments[0].AuthorIsUploader {
t.Fatalf("unexpected comments: %+v", comments)
}
if comments[0].TimeText != "2023-11-14 22:13" {
t.Errorf("timestamp fallback = %q, want UTC time", comments[0].TimeText)
}
if comments[1].Author != "Bob" || comments[1].TimeText != "yesterday" {
t.Fatalf("fallback comment fields not parsed: %+v", comments[1])
}
if len(heatmap) != 2 || heatmap[0].Width != 50 || heatmap[1].Height != 100 {
t.Fatalf("unexpected heatmap: %+v", heatmap)
}
}
func TestGetEngagementFiltersAndSortsHeatmap(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "heatmap-edge", `name = "Heatmap edge"`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"heatmap":[{"start_time":10,"end_time":20,"value":0.5},{"start_time":2,"end_time":4,"value":1},{"start_time":5,"end_time":4,"value":1},{"start_time":20,"end_time":21,"value":-1}]}`,
})
_, heatmap, err := svc.GetEngagement("heatmap-edge")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(heatmap) != 2 {
t.Fatalf("valid heatmap segment count = %d, want 2: %+v", len(heatmap), heatmap)
}
if heatmap[0].StartTime != 2 || heatmap[1].StartTime != 10 {
t.Errorf("heatmap order = %v, want ascending start time", heatmap)
}
for _, segment := range heatmap {
if segment.Width <= 0 || segment.Height < 8 || segment.Height > 100 {
t.Errorf("invalid normalized heatmap segment: %+v", segment)
}
}
}
func TestOrderCommentsNestsRepliesAndCapsDepth(t *testing.T) {
comments := []models.Comment{
{ID: "root", Parent: "root", Author: "root"},
{ID: "reply", Parent: "root", Author: "reply"},
{ID: "deep-1", Parent: "reply", Author: "deep-1"},
{ID: "deep-2", Parent: "deep-1", Author: "deep-2"},
{ID: "deep-3", Parent: "deep-2", Author: "deep-3"},
{ID: "deep-4", Parent: "deep-3", Author: "deep-4"},
{ID: "deep-5", Parent: "deep-4", Author: "deep-5"},
}
ordered := orderComments(comments)
if len(ordered) != len(comments) {
t.Fatalf("orderComments dropped comments: %d of %d", len(ordered), len(comments))
}
for i, want := range []string{"root", "reply", "deep-1", "deep-2", "deep-3", "deep-4", "deep-5"} {
if ordered[i].ID != want {
t.Errorf("comment %d = %q, want %q", i, ordered[i].ID, want)
}
}
if ordered[1].ReplyTo != "root" || ordered[len(ordered)-1].ReplyTo != "deep-4" {
t.Errorf("reply targets = %q, %q; want root and deep-4", ordered[1].ReplyTo, ordered[len(ordered)-1].ReplyTo)
}
if ordered[len(ordered)-1].Depth != 4 {
t.Errorf("deep reply depth = %d, want capped depth 4", ordered[len(ordered)-1].Depth)
}
}
func TestOrderCommentsHandlesMissingParentsCyclesAndSiblings(t *testing.T) {
comments := []models.Comment{
{ID: "root", Parent: "root", Author: "root"},
{ID: "second", Parent: " root ", Author: "second"},
{ID: "first", Parent: "root", Author: "first"},
{ID: "orphan", Parent: "missing", Author: "orphan"},
{ID: "cycle-a", Parent: "cycle-b", Author: "a"},
{ID: "cycle-b", Parent: "cycle-a", Author: "b"},
}
ordered := orderComments(comments)
if len(ordered) != len(comments) {
t.Fatalf("orderComments dropped malformed-tree comments: %d of %d", len(ordered), len(comments))
}
var ids []string
for _, comment := range ordered {
ids = append(ids, comment.ID)
}
if !reflect.DeepEqual(ids[:3], []string{"root", "second", "first"}) {
t.Errorf("sibling/root order = %v, want root then source-order siblings", ids)
}
for _, comment := range ordered {
if comment.ID == "orphan" && comment.Depth != 0 {
t.Errorf("orphan depth = %d, want top-level", comment.Depth)
}
}
}
func TestGetEngagementMalformedSidecarReturnsError(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "bad-engagement", `name = "Bad engagement"`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"comments":[`,
})
if _, _, err := svc.GetEngagement("bad-engagement"); err == nil {
t.Fatal("expected malformed engagement JSON to return an error")
}
}
func TestGetEngagementMissingSidecarIsEmpty(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "no-engagement", `name = "No engagement"`, map[string]string{"video.mp4": "dummy"})
comments, heatmap, err := svc.GetEngagement("no-engagement")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(comments) != 0 || len(heatmap) != 0 {
t.Fatalf("expected empty engagement, got comments=%v heatmap=%v", comments, heatmap)
}
}
Ainternal/service/formats.go
@@ -0,0 +1,109 @@
package service
import (
"bytes"
"encoding/json"
"fmt"
"os/exec"
"strconv"
"vidarchive/internal/models"
"vidarchive/internal/util"
)
func (s *DownloadService) ListFormats(url string) ([]*models.FormatInfo, error) {
// Use machine-readable JSON (-J) rather than scraping the human "-F" table,
// whose columns/separators shift between yt-dlp versions. stderr is captured
// separately so warnings can't corrupt the JSON on stdout.
cmd := exec.Command(s.cfg.YTDLPPath, "-J", "--no-warnings", url)
var stderr bytes.Buffer
cmd.Stderr = &stderr
output, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("yt-dlp -J failed: %w\n%s", err, stderr.String())
}
return parseFormatJSON(output)
}
// ytFormat mirrors the subset of yt-dlp's per-format JSON (-J) we surface.
// Numeric fields are pointers so an absent value (null/omitted) is distinct
// from a real zero.
type ytFormat struct {
FormatID string `json:"format_id"`
Ext string `json:"ext"`
Resolution string `json:"resolution"`
Width *int `json:"width"`
Height *int `json:"height"`
FPS *float64 `json:"fps"`
VCodec string `json:"vcodec"`
ACodec string `json:"acodec"`
AudioChannels *int `json:"audio_channels"`
Filesize *int64 `json:"filesize"`
FilesizeApprox *int64 `json:"filesize_approx"`
FormatNote string `json:"format_note"`
}
// parseFormatJSON reads yt-dlp's single-JSON dump (-J) and returns the available
// formats. For a single video the formats live at the top level; for a playlist
// URL we fall back to the first entry's formats so the picker still shows
// something useful.
func parseFormatJSON(data []byte) ([]*models.FormatInfo, error) {
var top struct {
Formats []ytFormat `json:"formats"`
Entries []struct {
Formats []ytFormat `json:"formats"`
} `json:"entries"`
}
if err := json.Unmarshal(data, &top); err != nil {
return nil, fmt.Errorf("parse yt-dlp JSON: %w", err)
}
raw := top.Formats
if len(raw) == 0 && len(top.Entries) > 0 {
raw = top.Entries[0].Formats
}
formats := make([]*models.FormatInfo, 0, len(raw))
for _, f := range raw {
formats = append(formats, f.toFormatInfo())
}
return formats, nil
}
func (f ytFormat) toFormatInfo() *models.FormatInfo {
fi := &models.FormatInfo{
ID: f.FormatID,
Ext: f.Ext,
Note: f.FormatNote,
}
switch {
case f.Resolution != "":
fi.Resolution = f.Resolution
case f.Width != nil && f.Height != nil && *f.Width > 0 && *f.Height > 0:
fi.Resolution = fmt.Sprintf("%dx%d", *f.Width, *f.Height)
}
if f.FPS != nil && *f.FPS > 0 {
fi.FPS = strconv.FormatFloat(*f.FPS, 'f', -1, 64)
}
if f.AudioChannels != nil && *f.AudioChannels > 0 {
fi.Channels = strconv.Itoa(*f.AudioChannels)
}
// Prefer the video codec; fall back to the audio codec for audio-only formats.
if f.VCodec != "" && f.VCodec != "none" {
fi.Codec = f.VCodec
} else if f.ACodec != "" && f.ACodec != "none" {
fi.Codec = f.ACodec
}
if f.Filesize != nil && *f.Filesize > 0 {
fi.FileSize = util.FormatBytes(*f.Filesize)
} else if f.FilesizeApprox != nil && *f.FilesizeApprox > 0 {
fi.FileSize = "~" + util.FormatBytes(*f.FilesizeApprox)
}
return fi
}
Ainternal/service/formats_test.go
@@ -0,0 +1,55 @@
package service
import (
"testing"
)
func TestParseFormatJSON(t *testing.T) {
data := []byte(`{
"id": "vid",
"formats": [
{"format_id": "18", "ext": "mp4", "resolution": "640x360", "fps": 30, "vcodec": "avc1", "acodec": "mp4a", "format_note": "360p"},
{"format_id": "137", "ext": "mp4", "width": 1920, "height": 1080, "fps": 60, "vcodec": "avc1", "acodec": "none", "filesize": 1048576, "format_note": "1080p"},
{"format_id": "233", "ext": "m4a", "resolution": "audio only", "vcodec": "none", "acodec": "mp4a", "audio_channels": 2, "format_note": "audio"}
]
}`)
formats, err := parseFormatJSON(data)
if err != nil {
t.Fatalf("parseFormatJSON: %v", err)
}
if len(formats) != 3 {
t.Fatalf("expected 3 formats, got %d: %+v", len(formats), formats)
}
if formats[0].ID != "18" || formats[0].Ext != "mp4" || formats[0].Resolution != "640x360" {
t.Errorf("format[0] = %+v", formats[0])
}
if formats[0].FPS != "30" {
t.Errorf("format[0] fps = %q, want 30", formats[0].FPS)
}
// Resolution is derived from width/height when no resolution string is present.
if formats[1].Resolution != "1920x1080" {
t.Errorf("format[1] resolution = %q, want 1920x1080", formats[1].Resolution)
}
if formats[1].FileSize != "1.0 MiB" {
t.Errorf("format[1] filesize = %q, want 1.0 MiB", formats[1].FileSize)
}
// Audio-only format: codec falls back to acodec and channels are populated.
if formats[2].Codec != "mp4a" {
t.Errorf("format[2] codec = %q, want mp4a", formats[2].Codec)
}
if formats[2].Channels != "2" {
t.Errorf("format[2] channels = %q, want 2", formats[2].Channels)
}
}
func TestParseFormatJSONPlaylistFallback(t *testing.T) {
// A playlist dump exposes formats under the first entry, not at the top level.
data := []byte(`{"_type":"playlist","entries":[{"id":"a","formats":[{"format_id":"18","ext":"mp4"}]}]}`)
formats, err := parseFormatJSON(data)
if err != nil {
t.Fatalf("parseFormatJSON: %v", err)
}
if len(formats) != 1 || formats[0].ID != "18" {
t.Fatalf("expected 1 format from entry fallback, got %+v", formats)
}
}
Ainternal/service/import.go
@@ -0,0 +1,374 @@
package service
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"os"
"os/exec"
"path/filepath"
"sort"
"strconv"
"strings"
"syscall"
"github.com/gabriel-vasile/mimetype"
"vidarchive/internal/models"
)
// resolveBaseLibraryDir returns the absolute library directory a download writes
// into, applying the optional per-download OutputDir while rejecting any path
// that escapes the library root.
func (s *DownloadService) resolveBaseLibraryDir(d *models.Download) (string, error) {
if !d.OutputDir.Valid || d.OutputDir.String == "" {
return s.cfg.LibraryDir, nil
}
// Reuse the library service's guard so both entry points enforce the boundary
// the same way — it resolves symlinks, which a plain prefix check does not.
dir, err := s.librarySvc.ResolveWithinLibrary(d.OutputDir.String)
if err != nil {
return "", fmt.Errorf("invalid output directory: %w", err)
}
return dir, nil
}
// importDownloadedItems moves each downloaded item from the temp dir into the
// library and returns the number of items successfully imported. Per-item
// failures are logged and skipped (a playlist with a few bad entries still
// imports the rest); a non-nil error means the import couldn't even start.
//
// A cancel stops the import between items and returns ctx.Err() with the count
// imported so far. Importing a long playlist takes real time (a move plus an
// ffprobe per file), so a deleted download must not keep filling the library.
func (s *DownloadService) importDownloadedItems(ctx context.Context, d *models.Download, tempDownloadDir, mode, ytdlpFlags string) (int, error) {
entries, err := os.ReadDir(tempDownloadDir)
if err != nil {
return 0, err
}
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
return 0, err
}
if err := os.MkdirAll(baseLibraryDir, 0755); err != nil {
return 0, err
}
var itemDirs []string
for _, entry := range entries {
if !entry.IsDir() {
continue
}
name := entry.Name()
if strings.HasPrefix(name, "item-") {
itemDirs = append(itemDirs, filepath.Join(tempDownloadDir, name))
}
}
sort.Strings(itemDirs)
imported := 0
for _, itemDir := range itemDirs {
if err := ctx.Err(); err != nil {
return imported, err
}
if err := s.importItemDir(ctx, d.URL, itemDir, baseLibraryDir, mode, ytdlpFlags); err != nil {
// A cancelled item isn't a bad item: stop instead of logging a warning
// for it and every one that follows.
if ctx.Err() != nil {
return imported, ctx.Err()
}
log.Printf("warning: failed to import item %s: %v", itemDir, err)
continue
}
imported++
}
// The temp dir (and any leftovers from failed imports) is removed by the
// caller's deferred cleanup, so partial state never leaks even on a crash.
return imported, nil
}
func (s *DownloadService) importItemDir(ctx context.Context, url, itemDir, baseLibraryDir, mode, ytdlpFlags string) error {
entries, err := os.ReadDir(itemDir)
if err != nil {
return err
}
var mediaFiles []os.DirEntry
var infoJSONPath string
var subtitleFiles []string
for _, entry := range entries {
if entry.IsDir() {
continue
}
name := entry.Name()
path := filepath.Join(itemDir, name)
ext := strings.ToLower(filepath.Ext(name))
if isInfoJSON(name) {
infoJSONPath = path
continue
}
if ext == ".vtt" || ext == ".srt" || ext == ".ass" || ext == ".ssa" {
subtitleFiles = append(subtitleFiles, path)
continue
}
mtype, err := mimetype.DetectFile(path)
if err == nil && mtype != nil && (strings.HasPrefix(mtype.String(), "audio/") || strings.HasPrefix(mtype.String(), "video/")) {
mediaFiles = append(mediaFiles, entry)
}
}
if len(mediaFiles) == 0 {
return fmt.Errorf("no media files found in %s", itemDir)
}
info := readInfoJSON(infoJSONPath)
name := s.deriveItemName(itemDir, info, mediaFiles)
videoID := info.ID
// Last point at which nothing has been written to the library yet: give up
// here on a cancel rather than part-way through, which would leave a folder
// with some of its files and no marker — or, in overwrite mode, delete the
// existing item and not replace it.
if err := ctx.Err(); err != nil {
return err
}
// Overwrite mode: replace the existing copy of this video in place rather than
// creating a duplicate folder. Removing the old dir lets uniqueDir reuse its
// name (or land on the new title if it changed upstream).
if mode == "overwrite" && videoID != "" {
if existing, ok := s.librarySvc.FindByVideoID(baseLibraryDir, videoID); ok {
if rel, err := filepath.Rel(s.cfg.LibraryDir, existing); err == nil {
s.librarySvc.evictCachedScan(filepath.ToSlash(rel))
}
os.RemoveAll(existing)
}
}
targetDir := s.uniqueDir(baseLibraryDir, name)
if err := os.MkdirAll(targetDir, 0755); err != nil {
return err
}
if infoJSONPath != "" {
if err := moveFile(infoJSONPath, filepath.Join(targetDir, "info.json")); err != nil {
return err
}
}
for _, entry := range mediaFiles {
if err := moveFile(filepath.Join(itemDir, entry.Name()), filepath.Join(targetDir, entry.Name())); err != nil {
return err
}
}
// Probe each media file's duration once, here in the worker (off the request
// path), and cache it in the marker so the library never has to probe while
// serving pages. Files we can't probe simply get no duration.
fileDurations := make(map[string]int)
for _, entry := range mediaFiles {
if d, ok := probeDuration(ctx, s.cfg.FFprobePath, filepath.Join(targetDir, entry.Name())); ok {
fileDurations[entry.Name()] = d
}
}
if len(subtitleFiles) > 0 {
subtitlesDir := filepath.Join(targetDir, subtitlesDirName)
if err := os.MkdirAll(subtitlesDir, 0755); err != nil {
return err
}
for _, sf := range subtitleFiles {
if err := moveFile(sf, filepath.Join(subtitlesDir, filepath.Base(sf))); err != nil {
return err
}
}
}
metadata := models.ItemMetadata{
Name: name,
SourceURL: url,
VideoID: videoID,
YtdlpFlags: ytdlpFlags,
FileDurations: fileDurations,
}
return s.librarySvc.writeMetadata(targetDir, metadata)
}
// infoJSON is the subset of yt-dlp's info.json VidArchive reads. ID is the
// stable item identity; within a single subscription's own directory it is
// enough to match items, so the extractor is not needed.
type infoJSON struct {
ID string `json:"id"`
Title string `json:"title"`
Description string `json:"description"`
WebpageURL string `json:"webpage_url"`
}
// readInfoJSON parses an info.json. A missing, unreadable or malformed file
// yields a zero-value struct: every caller treats absent fields as "unknown"
// and falls back, so there is nothing to distinguish.
func readInfoJSON(infoJSONPath string) infoJSON {
var info infoJSON
if infoJSONPath == "" {
return info
}
data, err := os.ReadFile(infoJSONPath)
if err != nil {
return info
}
if err := json.Unmarshal(data, &info); err != nil {
log.Printf("ignoring malformed %s: %v", infoJSONPath, err)
return infoJSON{}
}
return info
}
// isInfoJSON reports whether a file name is yt-dlp's metadata sidecar. yt-dlp
// writes "<title>.info.json" next to the media, but a bare "info.json" is what
// an already-imported item holds.
func isInfoJSON(name string) bool {
return name == "info.json" || strings.HasSuffix(name, ".info.json")
}
// findInfoJSON returns the path to an info.json directly inside itemDir, or "".
func findInfoJSON(itemDir string) string {
entries, err := os.ReadDir(itemDir)
if err != nil {
return ""
}
for _, entry := range entries {
if entry.IsDir() {
continue
}
name := entry.Name()
if isInfoJSON(name) {
return filepath.Join(itemDir, name)
}
}
return ""
}
// deriveItemName names the imported item after its title, falling back to the
// largest media file's base name when there is no usable info.json.
func (s *DownloadService) deriveItemName(itemDir string, info infoJSON, mediaFiles []os.DirEntry) string {
if info.Title != "" {
return sanitizeDirName(info.Title)
}
var largest os.DirEntry
var maxSize int64
for _, f := range mediaFiles {
st, err := os.Stat(filepath.Join(itemDir, f.Name()))
if err == nil && (largest == nil || st.Size() > maxSize) {
largest, maxSize = f, st.Size()
}
}
if largest == nil {
largest = mediaFiles[0]
}
base := strings.TrimSuffix(largest.Name(), filepath.Ext(largest.Name()))
return sanitizeDirName(base)
}
func (s *DownloadService) uniqueDir(base, name string) string {
dir := filepath.Join(base, name)
if _, err := os.Stat(dir); os.IsNotExist(err) {
return dir
}
for i := 1; ; i++ {
candidate := fmt.Sprintf("%s-%d", dir, i)
if _, err := os.Stat(candidate); os.IsNotExist(err) {
return candidate
}
}
}
// moveFile moves src to dst, falling back to copy-and-delete when the two are on
// different filesystems. The temp and library directories are independently
// configurable, so they can legitimately live on separate mounts — where a plain
// rename fails with EXDEV.
func moveFile(src, dst string) error {
if err := os.Rename(src, dst); err == nil {
return nil
} else if !errors.Is(err, syscall.EXDEV) {
return err
}
in, err := os.Open(src)
if err != nil {
return err
}
defer in.Close()
info, err := in.Stat()
if err != nil {
return err
}
out, err := os.OpenFile(dst, os.O_WRONLY|os.O_CREATE|os.O_TRUNC, info.Mode())
if err != nil {
return err
}
if _, err := io.Copy(out, in); err != nil {
out.Close()
os.Remove(dst)
return err
}
// Close explicitly: a deferred close would hide a flush error on the copy.
if err := out.Close(); err != nil {
os.Remove(dst)
return err
}
return os.Remove(src)
}
func sanitizeDirName(name string) string {
name = strings.TrimSpace(name)
replacer := strings.NewReplacer(
"/", "-",
"\\", "-",
":", "-",
"*", "-",
"?", "-",
"\"", "-",
"<", "-",
">", "-",
"|", "-",
)
name = replacer.Replace(name)
name = strings.TrimSpace(name)
// A name made only of dots resolves to the parent ("..") or to the target
// directory itself ("."), so joining it would place the item outside the
// library. Titles come from remote metadata, so refuse them here.
if strings.Trim(name, ".") == "" {
name = "untitled"
}
return name
}
// probeDuration returns the duration of a media file in whole seconds. The bool
// is false when ffprobe is unavailable or the file has no usable duration.
func probeDuration(ctx context.Context, ffprobePath, path string) (int, bool) {
out, err := exec.CommandContext(ctx, ffprobePath, "-v", "error",
"-show_entries", "format=duration",
"-of", "default=nw=1:nk=1", path).Output()
if err != nil {
return 0, false
}
f, err := strconv.ParseFloat(strings.TrimSpace(string(out)), 64)
if err != nil || f <= 0 {
return 0, false
}
return int(f + 0.5), true
}
Rinternal/service/import_test.go← internal/service/download_test.go
@@ -3,7 +3,6 @@ package service
import (
"context"
"database/sql"
"encoding/json"
"errors"
"fmt"
"os"
@@ -24,6 +23,13 @@ func TestSanitizeDirName(t *testing.T) {
{" trimmed ", "trimmed"},
{"", "untitled"},
{"///", "---"},
// A dot-only name would resolve to the target directory or its parent, so
// a title like this must not become a directory name.
{".", "untitled"},
{"..", "untitled"},
{" .. ", "untitled"},
{"...", "untitled"},
{".hidden", ".hidden"},
}
for _, tc := range tests {
if got := sanitizeDirName(tc.in); got != tc.want {
@@ -100,56 +106,6 @@ func dirEntry(t *testing.T, dir, name string) os.DirEntry {
return nil
}
func TestParseFormatJSON(t *testing.T) {
data := []byte(`{
"id": "vid",
"formats": [
{"format_id": "18", "ext": "mp4", "resolution": "640x360", "fps": 30, "vcodec": "avc1", "acodec": "mp4a", "format_note": "360p"},
{"format_id": "137", "ext": "mp4", "width": 1920, "height": 1080, "fps": 60, "vcodec": "avc1", "acodec": "none", "filesize": 1048576, "format_note": "1080p"},
{"format_id": "233", "ext": "m4a", "resolution": "audio only", "vcodec": "none", "acodec": "mp4a", "audio_channels": 2, "format_note": "audio"}
]
}`)
formats, err := parseFormatJSON(data)
if err != nil {
t.Fatalf("parseFormatJSON: %v", err)
}
if len(formats) != 3 {
t.Fatalf("expected 3 formats, got %d: %+v", len(formats), formats)
}
if formats[0].ID != "18" || formats[0].Ext != "mp4" || formats[0].Resolution != "640x360" {
t.Errorf("format[0] = %+v", formats[0])
}
if formats[0].FPS != "30" {
t.Errorf("format[0] fps = %q, want 30", formats[0].FPS)
}
// Resolution is derived from width/height when no resolution string is present.
if formats[1].Resolution != "1920x1080" {
t.Errorf("format[1] resolution = %q, want 1920x1080", formats[1].Resolution)
}
if formats[1].FileSize != "1.0 MiB" {
t.Errorf("format[1] filesize = %q, want 1.0 MiB", formats[1].FileSize)
}
// Audio-only format: codec falls back to acodec and channels are populated.
if formats[2].Codec != "mp4a" {
t.Errorf("format[2] codec = %q, want mp4a", formats[2].Codec)
}
if formats[2].Channels != "2" {
t.Errorf("format[2] channels = %q, want 2", formats[2].Channels)
}
}
func TestParseFormatJSONPlaylistFallback(t *testing.T) {
// A playlist dump exposes formats under the first entry, not at the top level.
data := []byte(`{"_type":"playlist","entries":[{"id":"a","formats":[{"format_id":"18","ext":"mp4"}]}]}`)
formats, err := parseFormatJSON(data)
if err != nil {
t.Fatalf("parseFormatJSON: %v", err)
}
if len(formats) != 1 || formats[0].ID != "18" {
t.Fatalf("expected 1 format from entry fallback, got %+v", formats)
}
}
// TestImportItemDir exercises the full import: media + info.json + subtitles get
// sorted into a named item directory with a marker. Uses ffmpeg to produce real
// media so the mimetype-based classification in importItemDir matches.
@@ -200,127 +156,6 @@ func TestImportItemDir(t *testing.T) {
}
}
func TestApplyMetadataPreservesEngagementAndPermissions(t *testing.T) {
e := newExecEnv(t)
existing := writeItem(t, e.cfg.LibraryDir, "existing", `name = "Old"
video_id = "old"
`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"id":"old","title":"Old","comments":[{"id":"c1","text":"archived"}],"heatmap":[{"start_time":0,"end_time":1,"value":1}],"filesize":123}`,
})
source := filepath.Join(e.scratch, "refreshed.info.json")
if err := os.WriteFile(source, []byte(`{"id":"new","title":"New","description":"updated","comments":[]}`), 0644); err != nil {
t.Fatal(err)
}
if err := e.svc.applyMetadata(existing, infoJSON{ID: "new", Title: "New", Description: "updated"}, source); err != nil {
t.Fatalf("applyMetadata: %v", err)
}
data, err := os.ReadFile(filepath.Join(existing, "info.json"))
if err != nil {
t.Fatal(err)
}
var got map[string]json.RawMessage
if err := json.Unmarshal(data, &got); err != nil {
t.Fatalf("parse installed info JSON: %v", err)
}
if _, ok := got["comments"]; !ok {
t.Fatal("refreshed sidecar lost archived comments")
}
if _, ok := got["heatmap"]; !ok {
t.Fatal("refreshed sidecar lost archived heatmap")
}
if string(got["title"]) != `"New"` {
t.Errorf("title = %s, want New", got["title"])
}
if mode := fileMode(t, filepath.Join(existing, "info.json")); mode != 0644 {
t.Errorf("info.json mode = %o, want 0644", mode)
}
}
func TestApplyMetadataRollsBackMarkerWhenSidecarInstallFails(t *testing.T) {
e := newExecEnv(t)
existing := writeItem(t, e.cfg.LibraryDir, "rollback", `name = "Old"
video_id = "old"
`, map[string]string{"video.mp4": "dummy"})
markerPath := filepath.Join(existing, itemMarkerName)
before, err := os.ReadFile(markerPath)
if err != nil {
t.Fatal(err)
}
// A directory at the destination makes the final rename fail after the
// marker has been written, exercising the rollback path.
infoPath := filepath.Join(existing, "info.json")
if err := os.Mkdir(infoPath, 0755); err != nil {
t.Fatal(err)
}
source := filepath.Join(e.scratch, "rollback.info.json")
if err := os.WriteFile(source, []byte(`{"id":"new","title":"New"}`), 0644); err != nil {
t.Fatal(err)
}
if err := e.svc.applyMetadata(existing, infoJSON{ID: "new", Title: "New"}, source); err == nil {
t.Fatal("expected sidecar installation to fail")
}
after, err := os.ReadFile(markerPath)
if err != nil {
t.Fatal(err)
}
if string(after) != string(before) {
t.Errorf("marker changed after failed sidecar install:\nbefore: %s\nafter: %s", before, after)
}
}
func fileMode(t *testing.T, path string) os.FileMode {
t.Helper()
info, err := os.Stat(path)
if err != nil {
t.Fatal(err)
}
return info.Mode().Perm()
}
func TestCheckReservedFlags(t *testing.T) {
cases := []struct {
name string
flags string
subscription bool
wantErr bool
}{
{"empty", "", false, false},
{"harmless", "--no-playlist --write-thumbnail", false, false},
{"output short", "-o foo.mp4", false, true},
{"output long", "--output foo.mp4", false, true},
{"output equals form", "--output=foo.mp4", false, true},
{"paths short", "-P /tmp", false, true},
{"cookies", "--cookies x.txt", false, true},
{"cookies inverse", "--no-cookies", false, true},
// Subscription-only reserved flags pass for normal downloads...
{"skip-download non-sub", "--skip-download", false, false},
{"write-info-json non-sub", "--write-info-json", false, false},
// ...but are rejected for subscription runs (and their inverses).
{"skip-download sub", "--skip-download", true, true},
{"no-skip-download sub", "--no-skip-download", true, true},
{"write-info-json sub", "--write-info-json", true, true},
{"no-write-info-json sub", "--no-write-info-json", true, true},
{"download-archive sub", "--download-archive a.txt", true, true},
// The "--flag=value" spelling names the same option as "--flag value",
// for the subscription-only table as well as the base one.
{"download-archive equals form sub", "--download-archive=a.txt", true, true},
// Base reserved flags still apply to subscriptions.
{"output sub", "-o x", true, true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
err := checkReservedFlags(tc.flags, tc.subscription)
if tc.wantErr != (err != nil) {
t.Errorf("checkReservedFlags(%q, %v) error = %v, wantErr %v", tc.flags, tc.subscription, err, tc.wantErr)
}
})
}
}
func TestImportDownloadedItemsRejectsOutputTraversal(t *testing.T) {
libDir := t.TempDir()
svc := &DownloadService{
Minternal/service/library.go
@@ -4,21 +4,17 @@ import (
"encoding/json"
"fmt"
"log"
"math"
"os"
"os/exec"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/BurntSushi/toml"
"vidarchive/internal/models"
"vidarchive/internal/util"
)
const itemMarkerName = ".vidarchive-item.toml"
@@ -37,11 +33,6 @@ var imageExts = map[string]struct{}{
".webp": {}, ".jpg": {}, ".jpeg": {}, ".png": {}, ".gif": {}, ".bmp": {},
}
// maxConcurrentThumbnails caps how many ffmpeg extraction processes may run at
// once, so a freshly loaded library page (which fires one thumbnail request per
// visible item) cannot spawn an unbounded ffmpeg storm.
const maxConcurrentThumbnails = 3
type LibraryService struct {
libraryDir string
// ffmpegPath/ffprobePath are the binaries used for thumbnail extraction,
@@ -551,248 +542,6 @@ func (s *LibraryService) GetMediaFile(relPath, filename string) (string, error)
return "", fmt.Errorf("media file not found")
}
// ThumbnailForFile returns the thumbnail for a specific media file within an
// item, extracting it on demand if needed. The bool is false when no thumbnail
// is available (file not found, audio-only, or extraction failed) so the caller
// can serve an icon. An empty/unmatched filename yields false — every thumbnail
// is keyed to a specific media file.
func (s *LibraryService) ThumbnailForFile(relPath, filename string) (string, bool) {
if filename == "" {
return "", false
}
item, err := s.GetByRelPath(relPath)
if err != nil {
return "", false
}
for i := range item.MediaFiles {
if item.MediaFiles[i].Filename != filename {
continue
}
mf := item.MediaFiles[i]
if path, ok := s.findExistingThumbnail(mf.Filepath); ok {
return path, true
}
return s.ensureThumbnailForFile(mf)
}
return "", false
}
// ensureThumbnailForFile returns an existing thumbnail for mf or extracts one,
// serializing concurrent extraction of the same file via a per-path mutex. It
// re-checks the disk under the lock so that whichever caller wins the race does
// the work and the rest reuse the result. A prior in-process failure short-
// circuits to avoid re-running ffmpeg on every request (see thumbFailed).
func (s *LibraryService) ensureThumbnailForFile(mf models.MediaFile) (string, bool) {
actual, _ := s.thumbLocks.LoadOrStore(mf.Filepath, &sync.Mutex{})
lock := actual.(*sync.Mutex)
lock.Lock()
defer lock.Unlock()
if path, ok := s.findExistingThumbnail(mf.Filepath); ok {
return path, true
}
// Trust a prior failure for this run rather than re-running ffmpeg every
// request; a restart clears thumbFailed and retries. Checked after the disk
// so a thumbnail that appears later (e.g. added manually) still wins.
if _, failed := s.thumbFailed.Load(mf.Filepath); failed {
return "", false
}
s.thumbSem <- struct{}{}
cur := atomic.AddInt32(&s.extractInFlight, 1)
for {
max := atomic.LoadInt32(&s.extractMaxConcurrent)
if cur <= max || atomic.CompareAndSwapInt32(&s.extractMaxConcurrent, max, cur) {
break
}
}
defer func() {
atomic.AddInt32(&s.extractInFlight, -1)
<-s.thumbSem
}()
atomic.AddInt32(&s.extractAttempts, 1)
path, err := s.extractThumbnail(mf)
if err != nil {
log.Printf("thumbnail extraction failed for %s: %v", mf.Filepath, err)
s.thumbFailed.Store(mf.Filepath, struct{}{})
return "", false
}
return path, true
}
func (s *LibraryService) findExistingThumbnail(path string) (string, bool) {
ext := filepath.Ext(path)
base := strings.TrimSuffix(path, ext) + ".thumbnail"
for _, candidate := range []string{base + ".webp", base + ".jpg", base + ".jpeg", base + ".png"} {
if info, err := os.Stat(candidate); err == nil && info.Size() > 0 {
return candidate, true
}
}
return "", false
}
// findImageAttachment returns the ordinal (0-based among attachment streams) of
// the best image attachment in a container — e.g. the cover.jpg/cover.webp that
// yt-dlp embeds into MKV with --embed-thumbnail — or -1 if there is none. Such
// covers are attachment streams, not attached_pic video streams, so they must be
// dumped with -dump_attachment rather than mapped like a normal stream.
func (s *LibraryService) findImageAttachment(path string) int {
cmd := exec.Command(s.ffprobePath, "-v", "error", "-show_streams", "-of", "json", path)
output, err := cmd.Output()
if err != nil {
return -1
}
var probe struct {
Streams []struct {
CodecType string `json:"codec_type"`
Tags struct {
Mimetype string `json:"mimetype"`
Filename string `json:"filename"`
} `json:"tags"`
} `json:"streams"`
}
if err := json.Unmarshal(output, &probe); err != nil {
return -1
}
best, bestScore, attachmentIdx := -1, 0, 0
for _, stream := range probe.Streams {
if stream.CodecType != "attachment" {
continue
}
idx := attachmentIdx
attachmentIdx++
if !strings.HasPrefix(strings.ToLower(stream.Tags.Mimetype), "image/") {
continue
}
score := 20
switch name := strings.ToLower(stream.Tags.Filename); {
case strings.Contains(name, "cover"):
score = 100
case strings.Contains(name, "thumbnail"), strings.Contains(name, "thumb"):
score = 80
case strings.Contains(name, "poster"):
score = 60
case strings.Contains(name, "art"):
score = 40
}
if score > bestScore {
best, bestScore = idx, score
}
}
return best
}
// thumbAttempt is one ffmpeg invocation that may produce a thumbnail at out.
type thumbAttempt struct {
label string
out string
args []string
}
func (s *LibraryService) extractThumbnail(mf models.MediaFile) (string, error) {
ext := filepath.Ext(mf.Filepath)
base := strings.TrimSuffix(mf.Filepath, ext) + ".thumbnail"
webpPath := base + ".webp"
jpgPath := base + ".jpg"
var attempts []thumbAttempt
// Collect every attempt's error so a genuine failure surfaces all of them
// rather than only the last fallback's stderr.
var attemptErrs []string
// Prefer an embedded image attachment (e.g. yt-dlp's cover.webp/cover.jpg in
// MKV): dump its raw bytes, then transcode to a canonical WebP.
if idx := s.findImageAttachment(mf.Filepath); idx >= 0 {
if rawPath, err := s.dumpAttachment(mf.Filepath, idx); err != nil {
attemptErrs = append(attemptErrs, fmt.Sprintf("attachment-dump: %v", err))
} else {
defer os.Remove(rawPath)
attempts = append(attempts, thumbAttempt{"attachment-webp", webpPath,
[]string{"-i", rawPath, "-c:v", "libwebp"}})
}
}
// Next try an embedded cover art video stream (attached_pic), e.g. mp3/mp4,
// falling back to jpeg if libwebp or webp encoding fails.
embedded := []string{"-i", mf.Filepath, "-map", "0:v", "-map", "-0:V", "-vframes", "1"}
attempts = append(attempts,
thumbAttempt{"embedded-webp", webpPath, append(embedded, "-c:v", "libwebp")},
thumbAttempt{"embedded-jpg", jpgPath, append(embedded, "-q:v", "2")},
)
if !mf.IsAudio {
seekTime := "00:00:01"
if mf.Duration > 0 {
seekTime = util.FormatClock(mf.Duration / 2)
}
frame := []string{"-ss", seekTime, "-i", mf.Filepath, "-vframes", "1"}
attempts = append(attempts,
thumbAttempt{"frame-webp", webpPath, append(frame, "-c:v", "libwebp")},
thumbAttempt{"frame-jpg", jpgPath, append(frame, "-q:v", "2")},
)
}
for _, a := range attempts {
path, err := s.tryWriteThumbnail(a.out, a.args)
if err == nil {
return path, nil
}
attemptErrs = append(attemptErrs, fmt.Sprintf("%s: %v", a.label, err))
}
if mf.IsAudio {
return "", fmt.Errorf("audio file has no thumbnail")
}
return "", fmt.Errorf("all thumbnail extraction attempts failed:\n%s", strings.Join(attemptErrs, "\n"))
}
// dumpAttachment extracts the raw bytes of attachment stream idx (e.g. the
// cover.jpg/cover.webp yt-dlp embeds into MKV with --embed-thumbnail) into a
// temp file and returns its path. The caller removes the file.
func (s *LibraryService) dumpAttachment(path string, idx int) (string, error) {
raw, err := os.CreateTemp("", "vidarchive-attachment-*")
if err != nil {
return "", err
}
rawPath := raw.Name()
raw.Close()
dumpArgs := []string{
fmt.Sprintf("-dump_attachment:t:%d", idx), rawPath,
"-i", path, "-y", "-t", "0", "-f", "null", "-",
}
if out, err := exec.Command(s.ffmpegPath, dumpArgs...).CombinedOutput(); err != nil {
os.Remove(rawPath)
return "", fmt.Errorf("%v\n%s", err, out)
}
return rawPath, nil
}
// tryWriteThumbnail runs ffmpeg with args to produce outputPath. The temp file
// keeps the final extension so ffmpeg can infer the output muxer (it cannot for
// a bare ".tmp" suffix), then is atomically renamed into place.
func (s *LibraryService) tryWriteThumbnail(outputPath string, args []string) (string, error) {
outExt := filepath.Ext(outputPath)
tmpPath := strings.TrimSuffix(outputPath, outExt) + ".tmp" + outExt
os.Remove(tmpPath)
cmd := exec.Command(s.ffmpegPath, append(args, tmpPath)...)
if output, err := cmd.CombinedOutput(); err != nil {
os.Remove(tmpPath)
return "", fmt.Errorf("ffmpeg failed: %v\n%s", err, string(output))
}
if info, err := os.Stat(tmpPath); err != nil || info.Size() == 0 {
os.Remove(tmpPath)
return "", fmt.Errorf("ffmpeg produced empty output")
}
if err := os.Rename(tmpPath, outputPath); err != nil {
os.Remove(tmpPath)
return "", err
}
return outputPath, nil
}
func (s *LibraryService) Delete(relPath string) error {
itemDir, err := s.resolveItemDir(relPath)
if err != nil {
@@ -890,529 +639,3 @@ func (s *LibraryService) PruneToIDSet(baseDir string, keep map[string]bool) (int
return removed, err
}
func (s *LibraryService) SubtitleDir(relPath string) string {
itemDir, err := s.resolveItemDir(relPath)
if err != nil {
return ""
}
return filepath.Join(itemDir, subtitlesDirName)
}
// GetSubtitlePath returns the .vtt path for a language. It errors when the item
// can't be resolved: joining onto an empty dir would yield a bare relative name
// that the caller would then serve relative to the process working directory.
func (s *LibraryService) GetSubtitlePath(relPath, lang string) (string, error) {
dir := s.SubtitleDir(relPath)
if dir == "" {
return "", fmt.Errorf("resolve subtitle dir for %q", relPath)
}
return filepath.Join(dir, lang+".vtt"), nil
}
func (s *LibraryService) GetSubtitles(relPath string) ([]models.SubtitleTrack, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, err
}
cacheDir := s.SubtitleDir(relPath)
entries, err := os.ReadDir(cacheDir)
if err == nil && len(entries) > 0 {
var tracks []models.SubtitleTrack
for _, entry := range entries {
if entry.IsDir() || filepath.Ext(entry.Name()) != ".vtt" {
continue
}
lang := strings.TrimSuffix(entry.Name(), ".vtt")
tracks = append(tracks, models.SubtitleTrack{
Lang: lang,
Label: lang,
Src: fmt.Sprintf("/media/item/%s/subtitles/%s", util.URLEncodePath(relPath), lang),
})
}
return tracks, nil
}
for _, mf := range item.MediaFiles {
if mf.IsAudio {
continue
}
streams, err := s.extractSubtitleInfo(mf.Filepath)
if err != nil || len(streams) == 0 {
continue
}
if err := os.MkdirAll(cacheDir, 0755); err != nil {
return nil, err
}
var tracks []models.SubtitleTrack
for _, stream := range streams {
lang := stream.Lang
if lang == "" {
lang = fmt.Sprintf("track%d", stream.Index)
}
outPath := filepath.Join(cacheDir, lang+".vtt")
if err := s.extractSubtitleToVTT(mf.Filepath, outPath, stream.Index); err != nil {
continue
}
tracks = append(tracks, models.SubtitleTrack{
Lang: lang,
Label: stream.Label,
Src: fmt.Sprintf("/media/item/%s/subtitles/%s", util.URLEncodePath(relPath), lang),
})
}
return tracks, nil
}
return nil, nil
}
// GetEngagement reads optional comments and a playback heatmap from the item's
// yt-dlp info sidecar. These fields are intentionally not copied into the
// marker: the sidecar remains the source of truth and old items simply return
// empty data when the fields are absent.
func (s *LibraryService) GetEngagement(relPath string) ([]models.Comment, []models.HeatmapSegment, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, nil, err
}
_, infoPath, err := s.listItemFiles(item.DirPath)
if err != nil || infoPath == "" {
return nil, nil, err
}
data, err := os.ReadFile(infoPath)
if err != nil {
return nil, nil, err
}
var raw struct {
Comments []struct {
ID string `json:"id"`
Parent string `json:"parent"`
Author string `json:"author"`
Channel string `json:"channel"`
Text string `json:"text"`
Timestamp float64 `json:"timestamp"`
TimeText string `json:"time_text"`
LikeCount int `json:"like_count"`
AuthorIsUploader bool `json:"author_is_uploader"`
} `json:"comments"`
Heatmap []struct {
StartTime float64 `json:"start_time"`
EndTime float64 `json:"end_time"`
Value float64 `json:"value"`
} `json:"heatmap"`
}
if err := json.Unmarshal(data, &raw); err != nil {
return nil, nil, fmt.Errorf("parse engagement info JSON: %w", err)
}
comments := make([]models.Comment, 0, len(raw.Comments))
for _, c := range raw.Comments {
author := c.Author
if author == "" {
author = c.Channel
}
if c.Text == "" && author == "" {
continue
}
comment := models.Comment{
ID: c.ID,
Parent: c.Parent,
Author: author,
Text: c.Text,
Timestamp: int64(c.Timestamp),
TimeText: c.TimeText,
LikeCount: c.LikeCount,
AuthorIsUploader: c.AuthorIsUploader,
}
if comment.TimeText == "" && comment.Timestamp > 0 {
comment.TimeText = time.Unix(comment.Timestamp, 0).UTC().Format("2006-01-02 15:04")
}
comments = append(comments, comment)
}
comments = orderComments(comments)
maxValue := 0.0
for _, h := range raw.Heatmap {
if h.Value > maxValue && !math.IsNaN(h.Value) && !math.IsInf(h.Value, 0) {
maxValue = h.Value
}
}
endTime := 0.0
for _, h := range raw.Heatmap {
if h.EndTime > endTime {
endTime = h.EndTime
}
}
heatmap := make([]models.HeatmapSegment, 0, len(raw.Heatmap))
for _, h := range raw.Heatmap {
if h.EndTime <= h.StartTime || h.StartTime < 0 || h.Value < 0 ||
math.IsNaN(h.StartTime) || math.IsNaN(h.EndTime) || math.IsNaN(h.Value) ||
math.IsInf(h.StartTime, 0) || math.IsInf(h.EndTime, 0) || math.IsInf(h.Value, 0) {
continue
}
width := 0.0
if endTime > 0 {
width = (h.EndTime - h.StartTime) / endTime * 100
}
height := 8.0
if maxValue > 0 {
height += h.Value / maxValue * 92
}
heatmap = append(heatmap, models.HeatmapSegment{
StartTime: h.StartTime,
EndTime: h.EndTime,
Value: h.Value,
Width: width,
Height: height,
})
}
sort.SliceStable(heatmap, func(i, j int) bool {
return heatmap[i].StartTime < heatmap[j].StartTime
})
return comments, heatmap, nil
}
// orderComments groups replies beneath their parent while preserving the
// source order among siblings. Depth is capped for presentation so malformed
// or unusually deep reply chains cannot make the UI progressively narrower.
func orderComments(comments []models.Comment) []models.Comment {
if len(comments) < 2 {
return comments
}
byID := make(map[string]int, len(comments))
for i, comment := range comments {
if comment.ID != "" {
if _, exists := byID[comment.ID]; !exists {
byID[comment.ID] = i
}
}
}
children := make(map[int][]int)
var roots []int
for i, comment := range comments {
parent := strings.TrimSpace(comment.Parent)
if parent == "" || strings.EqualFold(parent, "root") {
roots = append(roots, i)
continue
}
parentIndex, ok := byID[parent]
if !ok || parentIndex == i {
roots = append(roots, i)
continue
}
children[parentIndex] = append(children[parentIndex], i)
}
ordered := make([]models.Comment, 0, len(comments))
visited := make([]bool, len(comments))
stack := make([]struct {
index int
depth int
}, 0, len(comments))
walk := func(index, depth int) {
stack = append(stack, struct {
index int
depth int
}{index, depth})
for len(stack) > 0 {
last := len(stack) - 1
entry := stack[last]
stack = stack[:last]
if visited[entry.index] {
continue
}
visited[entry.index] = true
comment := comments[entry.index]
if parentIndex, ok := byID[strings.TrimSpace(comment.Parent)]; ok && parentIndex != entry.index {
comment.ReplyTo = comments[parentIndex].Author
}
if entry.depth > 4 {
comment.Depth = 4
} else {
comment.Depth = entry.depth
}
ordered = append(ordered, comment)
childrenForComment := children[entry.index]
for i := len(childrenForComment) - 1; i >= 0; i-- {
stack = append(stack, struct {
index int
depth int
}{childrenForComment[i], entry.depth + 1})
}
}
}
for _, root := range roots {
walk(root, 0)
}
// Cycles or references to invalid parents are rendered as top-level comments
// rather than being dropped.
for i := range comments {
if !visited[i] {
walk(i, 0)
}
}
return ordered
}
type subtitleStream struct {
Index int
Lang string
Label string
}
func (s *LibraryService) extractSubtitleInfo(path string) ([]subtitleStream, error) {
cmd := exec.Command(s.ffprobePath,
"-v", "error",
"-show_streams",
"-select_streams", "s",
"-of", "json",
path,
)
output, err := cmd.Output()
if err != nil {
return nil, err
}
var probe struct {
Streams []struct {
Index int `json:"index"`
CodecName string `json:"codec_name"`
Tags struct {
Language string `json:"language"`
Title string `json:"title"`
} `json:"tags"`
} `json:"streams"`
}
if err := json.Unmarshal(output, &probe); err != nil {
return nil, err
}
var streams []subtitleStream
subIndex := 0
for _, stream := range probe.Streams {
switch stream.CodecName {
case "subrip", "ass", "ssa", "webvtt", "mov_text":
lang := stream.Tags.Language
if lang == "" {
lang = fmt.Sprintf("track%d", subIndex)
}
label := stream.Tags.Title
if label == "" {
label = strings.ToUpper(lang)
}
streams = append(streams, subtitleStream{
Index: subIndex,
Lang: lang,
Label: label,
})
subIndex++
}
}
return streams, nil
}
func (s *LibraryService) extractSubtitleToVTT(inputPath, outputPath string, streamIndex int) error {
cmd := exec.Command(s.ffmpegPath,
"-i", inputPath,
"-map", fmt.Sprintf("0:s:%d", streamIndex),
"-f", "webvtt",
outputPath,
"-y",
)
output, err := cmd.CombinedOutput()
if err != nil {
return fmt.Errorf("ffmpeg subtitle extraction failed: %w\nOutput: %s", err, string(output))
}
return nil
}
// GetMetadata probes media details for the named file within an item. An empty
// filename (or one that doesn't match) falls back to the item's primary media
// file, so the detail view shows metadata for whichever file is selected.
func (s *LibraryService) GetMetadata(relPath, filename string) (*MediaMetadata, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, err
}
var target *models.MediaFile
if filename != "" {
for i := range item.MediaFiles {
if item.MediaFiles[i].Filename == filename {
target = &item.MediaFiles[i]
break
}
}
}
if target == nil {
target = primaryMediaFile(item)
}
if target == nil {
return nil, fmt.Errorf("no media file")
}
return s.probeMedia(target.Filepath)
}
func (s *LibraryService) probeMedia(path string) (*MediaMetadata, error) {
info, err := os.Stat(path)
if err != nil {
return nil, err
}
meta := &MediaMetadata{
FileSize: info.Size(),
}
cmd := exec.Command(s.ffprobePath,
"-v", "error",
"-show_format",
"-show_streams",
"-of", "json",
path,
)
output, err := cmd.Output()
if err != nil {
return meta, nil
}
var probe ffprobeOutput
if err := json.Unmarshal(output, &probe); err != nil {
return meta, nil
}
if probe.Format.FormatName != "" {
parts := strings.Split(probe.Format.FormatName, ",")
meta.Container = parts[0]
}
for _, stream := range probe.Streams {
switch stream.CodecType {
case "video":
vs := VideoStream{
Codec: stream.CodecName,
Profile: stream.Profile,
Width: stream.Width,
Height: stream.Height,
FPS: parseFPS(stream.RFrameRate),
PixelFormat: stream.PixFmt,
Bitrate: formatBitrate(stream.BitRate),
}
meta.VideoStreams = append(meta.VideoStreams, vs)
if stream.Width > 0 && stream.Height > 0 {
meta.Resolution = fmt.Sprintf("%dx%d", stream.Width, stream.Height)
}
case "audio":
as := AudioStream{
Codec: stream.CodecName,
SampleRate: stream.SampleRate,
Channels: stream.Channels,
ChannelLayout: stream.ChannelLayout,
SampleFormat: stream.SampleFmt,
Bitrate: formatBitrate(stream.BitRate),
Language: stream.Tags.Language,
}
meta.AudioStreams = append(meta.AudioStreams, as)
case "subtitle":
ss := SubtitleStream{
Codec: stream.CodecName,
Language: stream.Tags.Language,
Title: stream.Tags.Title,
}
meta.SubtitleStreams = append(meta.SubtitleStreams, ss)
}
}
return meta, nil
}
func formatBitrate(bitRate string) string {
if bitRate == "" {
return ""
}
br, err := strconv.ParseInt(bitRate, 10, 64)
if err != nil {
return ""
}
return fmt.Sprintf("%d", br/1000)
}
func parseFPS(rate string) string {
var num, den float64
if _, err := fmt.Sscanf(rate, "%f/%f", &num, &den); err != nil || den == 0 {
return ""
}
fps := num / den
if fps == math.Trunc(fps) {
return fmt.Sprintf("%.0f", fps)
}
return fmt.Sprintf("%.2f", fps)
}
type MediaMetadata struct {
Container string
Resolution string
FileSize int64
VideoStreams []VideoStream
AudioStreams []AudioStream
SubtitleStreams []SubtitleStream
}
type VideoStream struct {
Codec string
Profile string
Width int
Height int
FPS string
PixelFormat string
Bitrate string
}
type AudioStream struct {
Codec string
SampleRate string
Channels int
ChannelLayout string
SampleFormat string
Bitrate string
Language string
}
type SubtitleStream struct {
Language string
Title string
Codec string
}
type ffprobeOutput struct {
Format struct {
FormatName string `json:"format_name"`
BitRate string `json:"bit_rate"`
} `json:"format"`
Streams []ffprobeStream `json:"streams"`
}
type ffprobeStream struct {
Index int `json:"index"`
CodecName string `json:"codec_name"`
CodecType string `json:"codec_type"`
Profile string `json:"profile"`
Width int `json:"width"`
Height int `json:"height"`
RFrameRate string `json:"r_frame_rate"`
AvgFrameRate string `json:"avg_frame_rate"`
PixFmt string `json:"pix_fmt"`
SampleRate string `json:"sample_rate"`
Channels int `json:"channels"`
ChannelLayout string `json:"channel_layout"`
SampleFmt string `json:"sample_fmt"`
BitRate string `json:"bit_rate"`
Tags struct {
Language string `json:"language"`
Title string `json:"title"`
} `json:"tags"`
}
Minternal/service/library_test.go
@@ -1,14 +1,10 @@
package service
import (
"fmt"
"os"
"os/exec"
"path/filepath"
"reflect"
"strings"
"sync"
"sync/atomic"
"testing"
"vidarchive/internal/models"
@@ -69,44 +65,6 @@ func makeTestVideoSize(t *testing.T, path, size string) {
}
}
// makeVideoWithCoverAttachment writes an MKV at outPath with a 64x64 video and
// an embedded "cover.webp" image attachment of the given size — mirroring how
// yt-dlp embeds thumbnails into MKV (a true attachment stream, not an
// attached_pic video stream; note ffmpeg muxes an attached .png as a video
// stream, so webp/jpg must be used to get a real attachment).
func makeVideoWithCoverAttachment(t *testing.T, outPath, coverSize string) {
t.Helper()
scratch := t.TempDir()
base := filepath.Join(scratch, "base.mp4")
makeTestVideoSize(t, base, "64x64")
cover := filepath.Join(scratch, "cover.webp")
if out, err := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-f", "lavfi", "-i", "color=red:size="+coverSize+":duration=1",
"-frames:v", "1", cover, "-y").CombinedOutput(); err != nil {
t.Fatalf("make cover: %v\n%s", err, out)
}
if out, err := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-i", base, "-attach", cover, "-metadata:s:t:0", "mimetype=image/webp",
"-c", "copy", outPath, "-y").CombinedOutput(); err != nil {
t.Fatalf("attach cover: %v\n%s", err, out)
}
}
// probeImageSize returns the pixel dimensions of an image/video file.
func probeImageSize(t *testing.T, path string) (int, int) {
t.Helper()
out, err := exec.Command("ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_entries", "stream=width,height", "-of", "csv=p=0:s=x", path).Output()
if err != nil {
t.Fatalf("probe %s: %v", path, err)
}
var w, h int
if _, err := fmt.Sscanf(strings.TrimSpace(string(out)), "%dx%d", &w, &h); err != nil {
t.Fatalf("parse dimensions %q: %v", out, err)
}
return w, h
}
// --- pure helper tests ---
func TestInfoDuration(t *testing.T) {
@@ -133,21 +91,6 @@ func TestInfoDuration(t *testing.T) {
}
}
func TestParseFPS(t *testing.T) {
tests := []struct{ in, want string }{
{"30/1", "30"},
{"30000/1001", "29.97"},
{"0/0", ""},
{"", ""},
{"garbage", ""},
}
for _, tc := range tests {
if got := parseFPS(tc.in); got != tc.want {
t.Errorf("parseFPS(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
func TestInfoString(t *testing.T) {
if _, ok := infoString(nil, "title"); ok {
t.Error("nil info should return false")
@@ -200,147 +143,6 @@ func TestPrimaryMediaFile(t *testing.T) {
}
}
func TestFormatBitrate(t *testing.T) {
tests := []struct{ in, want string }{
{"128000", "128"},
{"", ""},
{"notanumber", ""},
}
for _, tc := range tests {
if got := formatBitrate(tc.in); got != tc.want {
t.Errorf("formatBitrate(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
func TestGetEngagement(t *testing.T) {
svc, dir := newLibrary(t)
info := `{"comments":[{"author":"Alice","text":"hello <world>","timestamp":1700000000,"like_count":3,"author_is_uploader":true},{"channel":"Bob","text":"second","time_text":"yesterday"}],"heatmap":[{"start_time":0,"end_time":10,"value":0.5},{"start_time":10,"end_time":20,"value":1.0}]}`
writeItem(t, dir, "engagement", `name = "Engagement"`, map[string]string{
"video.mp4": "dummy",
"info.json": info,
})
comments, heatmap, err := svc.GetEngagement("engagement")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(comments) != 2 || comments[0].Author != "Alice" || comments[0].LikeCount != 3 || !comments[0].AuthorIsUploader {
t.Fatalf("unexpected comments: %+v", comments)
}
if comments[0].TimeText != "2023-11-14 22:13" {
t.Errorf("timestamp fallback = %q, want UTC time", comments[0].TimeText)
}
if comments[1].Author != "Bob" || comments[1].TimeText != "yesterday" {
t.Fatalf("fallback comment fields not parsed: %+v", comments[1])
}
if len(heatmap) != 2 || heatmap[0].Width != 50 || heatmap[1].Height != 100 {
t.Fatalf("unexpected heatmap: %+v", heatmap)
}
}
func TestGetEngagementFiltersAndSortsHeatmap(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "heatmap-edge", `name = "Heatmap edge"`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"heatmap":[{"start_time":10,"end_time":20,"value":0.5},{"start_time":2,"end_time":4,"value":1},{"start_time":5,"end_time":4,"value":1},{"start_time":20,"end_time":21,"value":-1}]}`,
})
_, heatmap, err := svc.GetEngagement("heatmap-edge")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(heatmap) != 2 {
t.Fatalf("valid heatmap segment count = %d, want 2: %+v", len(heatmap), heatmap)
}
if heatmap[0].StartTime != 2 || heatmap[1].StartTime != 10 {
t.Errorf("heatmap order = %v, want ascending start time", heatmap)
}
for _, segment := range heatmap {
if segment.Width <= 0 || segment.Height < 8 || segment.Height > 100 {
t.Errorf("invalid normalized heatmap segment: %+v", segment)
}
}
}
func TestOrderCommentsNestsRepliesAndCapsDepth(t *testing.T) {
comments := []models.Comment{
{ID: "root", Parent: "root", Author: "root"},
{ID: "reply", Parent: "root", Author: "reply"},
{ID: "deep-1", Parent: "reply", Author: "deep-1"},
{ID: "deep-2", Parent: "deep-1", Author: "deep-2"},
{ID: "deep-3", Parent: "deep-2", Author: "deep-3"},
{ID: "deep-4", Parent: "deep-3", Author: "deep-4"},
{ID: "deep-5", Parent: "deep-4", Author: "deep-5"},
}
ordered := orderComments(comments)
if len(ordered) != len(comments) {
t.Fatalf("orderComments dropped comments: %d of %d", len(ordered), len(comments))
}
for i, want := range []string{"root", "reply", "deep-1", "deep-2", "deep-3", "deep-4", "deep-5"} {
if ordered[i].ID != want {
t.Errorf("comment %d = %q, want %q", i, ordered[i].ID, want)
}
}
if ordered[1].ReplyTo != "root" || ordered[len(ordered)-1].ReplyTo != "deep-4" {
t.Errorf("reply targets = %q, %q; want root and deep-4", ordered[1].ReplyTo, ordered[len(ordered)-1].ReplyTo)
}
if ordered[len(ordered)-1].Depth != 4 {
t.Errorf("deep reply depth = %d, want capped depth 4", ordered[len(ordered)-1].Depth)
}
}
func TestOrderCommentsHandlesMissingParentsCyclesAndSiblings(t *testing.T) {
comments := []models.Comment{
{ID: "root", Parent: "root", Author: "root"},
{ID: "second", Parent: " root ", Author: "second"},
{ID: "first", Parent: "root", Author: "first"},
{ID: "orphan", Parent: "missing", Author: "orphan"},
{ID: "cycle-a", Parent: "cycle-b", Author: "a"},
{ID: "cycle-b", Parent: "cycle-a", Author: "b"},
}
ordered := orderComments(comments)
if len(ordered) != len(comments) {
t.Fatalf("orderComments dropped malformed-tree comments: %d of %d", len(ordered), len(comments))
}
var ids []string
for _, comment := range ordered {
ids = append(ids, comment.ID)
}
if !reflect.DeepEqual(ids[:3], []string{"root", "second", "first"}) {
t.Errorf("sibling/root order = %v, want root then source-order siblings", ids)
}
for _, comment := range ordered {
if comment.ID == "orphan" && comment.Depth != 0 {
t.Errorf("orphan depth = %d, want top-level", comment.Depth)
}
}
}
func TestGetEngagementMalformedSidecarReturnsError(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "bad-engagement", `name = "Bad engagement"`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"comments":[`,
})
if _, _, err := svc.GetEngagement("bad-engagement"); err == nil {
t.Fatal("expected malformed engagement JSON to return an error")
}
}
func TestGetEngagementMissingSidecarIsEmpty(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "no-engagement", `name = "No engagement"`, map[string]string{"video.mp4": "dummy"})
comments, heatmap, err := svc.GetEngagement("no-engagement")
if err != nil {
t.Fatalf("GetEngagement: %v", err)
}
if len(comments) != 0 || len(heatmap) != 0 {
t.Fatalf("expected empty engagement, got comments=%v heatmap=%v", comments, heatmap)
}
}
// --- listItemFiles / GetAll ---
func TestListItemFilesClassification(t *testing.T) {
@@ -672,263 +474,3 @@ func TestDeleteRejectsLibraryRoot(t *testing.T) {
}
// --- Thumbnail behavior ---
func TestThumbnailPrefersExistingGenerated(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "item", "name = \"I\"\nduration = -1\n", map[string]string{
"video.mp4": "v",
"video.thumbnail.webp": "GENERATED",
})
path, ok := svc.ThumbnailForFile("item", "video.mp4")
if !ok {
t.Fatal("expected a thumbnail")
}
if !strings.HasSuffix(path, "video.thumbnail.webp") {
t.Errorf("expected generated thumbnail, got %q", path)
}
}
func TestThumbnailExtractsOnDemand(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
path, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("expected on-demand extraction to succeed")
}
if info, err := os.Stat(path); err != nil || info.Size() == 0 {
t.Fatalf("thumbnail file missing/empty: %v", err)
}
if !strings.Contains(filepath.Base(path), ".thumbnail.") {
t.Errorf("unexpected thumbnail name %q", path)
}
// No leftover temp files from the atomic-write path.
entries, _ := os.ReadDir(itemDir)
for _, e := range entries {
if strings.Contains(e.Name(), ".tmp") {
t.Errorf("leftover temp file %q", e.Name())
}
}
}
func TestThumbnailRetriesAfterDeletion(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
first, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("first extraction failed")
}
if err := os.Remove(first); err != nil {
t.Fatal(err)
}
// A failure/absence must not be cached permanently: re-request re-extracts.
second, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("re-extraction after deletion failed (failure was cached)")
}
if info, err := os.Stat(second); err != nil || info.Size() == 0 {
t.Fatalf("re-extracted thumbnail missing/empty: %v", err)
}
}
func TestThumbnailConcurrentSingleExtraction(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
var wg sync.WaitGroup
for i := 0; i < 8; i++ {
wg.Add(1)
go func() {
defer wg.Done()
if _, ok := svc.ThumbnailForFile("vid", "vid.mp4"); !ok {
t.Error("concurrent Thumbnail failed")
}
}()
}
wg.Wait()
// Exactly one generated thumbnail, no temp leftovers despite the race.
entries, _ := os.ReadDir(itemDir)
var thumbs int
for _, e := range entries {
if strings.Contains(e.Name(), ".thumbnail.") {
thumbs++
}
if strings.Contains(e.Name(), ".tmp") {
t.Errorf("leftover temp file %q", e.Name())
}
}
if thumbs != 1 {
t.Errorf("expected exactly 1 generated thumbnail, got %d", thumbs)
}
}
func TestThumbnailForFileIsPerFile(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "multi", "name = \"M\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "a.mp4"))
makeTestVideo(t, filepath.Join(itemDir, "b.mp4"))
pa, ok := svc.ThumbnailForFile("multi", "a.mp4")
if !ok {
t.Fatal("thumbnail for a.mp4 failed")
}
pb, ok := svc.ThumbnailForFile("multi", "b.mp4")
if !ok {
t.Fatal("thumbnail for b.mp4 failed")
}
if pa == pb {
t.Errorf("expected distinct per-file thumbnails, both = %q", pa)
}
if !strings.Contains(filepath.Base(pa), "a.thumbnail.") {
t.Errorf("a.mp4 thumbnail name = %q", filepath.Base(pa))
}
if !strings.Contains(filepath.Base(pb), "b.thumbnail.") {
t.Errorf("b.mp4 thumbnail name = %q", filepath.Base(pb))
}
// An unknown file yields no thumbnail (caller falls back to an icon).
if _, ok := svc.ThumbnailForFile("multi", "nope.mp4"); ok {
t.Error("unknown file should not produce a thumbnail")
}
// Every thumbnail is keyed to a specific file: an empty filename yields none.
if _, ok := svc.ThumbnailForFile("multi", ""); ok {
t.Error("empty filename should not produce a thumbnail")
}
}
func TestGetMetadataForSelectedFile(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "multi", "name = \"M\"\nduration = -1\n", nil)
makeTestVideoSize(t, filepath.Join(itemDir, "small.mp4"), "64x64")
makeTestVideoSize(t, filepath.Join(itemDir, "big.mp4"), "128x72")
m1, err := svc.GetMetadata("multi", "small.mp4")
if err != nil {
t.Fatalf("GetMetadata small: %v", err)
}
if m1.Resolution != "64x64" {
t.Errorf("small.mp4 resolution = %q, want 64x64", m1.Resolution)
}
m2, err := svc.GetMetadata("multi", "big.mp4")
if err != nil {
t.Fatalf("GetMetadata big: %v", err)
}
if m2.Resolution != "128x72" {
t.Errorf("big.mp4 resolution = %q, want 128x72", m2.Resolution)
}
}
func TestThumbnailConcurrencyBounded(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "many", "name = \"M\"\nduration = -1\n", nil)
const n = 8
for i := 0; i < n; i++ {
makeTestVideo(t, filepath.Join(itemDir, fmt.Sprintf("c%d.mp4", i)))
}
var wg sync.WaitGroup
for i := 0; i < n; i++ {
i := i
wg.Add(1)
go func() {
defer wg.Done()
svc.ThumbnailForFile("many", fmt.Sprintf("c%d.mp4", i))
}()
}
wg.Wait()
max := atomic.LoadInt32(&svc.extractMaxConcurrent)
if max > maxConcurrentThumbnails {
t.Errorf("peak concurrent extractions %d exceeded cap %d", max, maxConcurrentThumbnails)
}
if max < 1 {
t.Error("expected at least one extraction to run")
}
}
func TestThumbnailUsesEmbeddedAttachment(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "att", "name = \"A\"\nduration = -1\n", nil)
mkv := filepath.Join(itemDir, "v.mkv")
// 100x100 cover so it's distinguishable from a 64x64 video frame.
makeVideoWithCoverAttachment(t, mkv, "100x100")
if idx := svc.findImageAttachment(mkv); idx < 0 {
t.Fatal("findImageAttachment did not find the embedded cover")
}
path, ok := svc.ThumbnailForFile("att", "v.mkv")
if !ok {
t.Fatal("thumbnail extraction failed")
}
// The embedded cover (100x100) must be used in preference to a video frame
// (which would be 64x64) — this is the regression the refactor introduced.
if w, h := probeImageSize(t, path); w != 100 || h != 100 {
t.Errorf("thumbnail is %dx%d, expected 100x100 from the embedded cover (got a video frame instead)", w, h)
}
}
func TestThumbnailNegativeCacheSkipsReextraction(t *testing.T) {
svc, dir := newLibrary(t)
// A file ffmpeg cannot extract a thumbnail from: every attempt fails.
writeItem(t, dir, "bad", "name = \"B\"\nduration = -1\n", map[string]string{
"broken.mp4": "not actually a video",
})
if _, ok := svc.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("expected extraction to fail for a non-video file")
}
attempts1 := atomic.LoadInt32(&svc.extractAttempts)
if attempts1 == 0 {
t.Fatal("expected at least one extraction attempt")
}
// A second request is served from the in-process negative cache: no new
// ffmpeg attempt.
if _, ok := svc.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("expected the cached failure to persist")
}
if attempts2 := atomic.LoadInt32(&svc.extractAttempts); attempts2 != attempts1 {
t.Errorf("negative cache should prevent re-extraction; attempts %d -> %d", attempts1, attempts2)
}
// The cache is in-process only: a fresh service (≈ a restart) retries.
fresh := NewLibraryService(dir, "ffmpeg", "ffprobe")
if _, ok := fresh.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("fresh service still fails (file is unextractable)")
}
if atomic.LoadInt32(&fresh.extractAttempts) == 0 {
t.Error("a fresh service should retry extraction, not inherit the negative cache")
}
}
func TestThumbnailAudioOnlyHasNone(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "aud", "name = \"A\"\nduration = -1\n", nil)
// A real audio file with no cover art.
cmd := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-f", "lavfi", "-i", "sine=frequency=440:duration=1",
filepath.Join(itemDir, "aud.mp3"), "-y")
if out, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("make audio: %v\n%s", err, out)
}
if path, ok := svc.ThumbnailForFile("aud", "aud.mp3"); ok {
t.Errorf("audio-only item should have no thumbnail, got %q", path)
}
}
Ainternal/service/media_probe.go
@@ -0,0 +1,200 @@
package service
import (
"encoding/json"
"fmt"
"math"
"os"
"os/exec"
"strconv"
"strings"
"vidarchive/internal/models"
)
// GetMetadata probes media details for the named file within an item. An empty
// filename (or one that doesn't match) falls back to the item's primary media
// file, so the detail view shows metadata for whichever file is selected.
func (s *LibraryService) GetMetadata(relPath, filename string) (*MediaMetadata, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, err
}
var target *models.MediaFile
if filename != "" {
for i := range item.MediaFiles {
if item.MediaFiles[i].Filename == filename {
target = &item.MediaFiles[i]
break
}
}
}
if target == nil {
target = primaryMediaFile(item)
}
if target == nil {
return nil, fmt.Errorf("no media file")
}
return s.probeMedia(target.Filepath)
}
func (s *LibraryService) probeMedia(path string) (*MediaMetadata, error) {
info, err := os.Stat(path)
if err != nil {
return nil, err
}
meta := &MediaMetadata{
FileSize: info.Size(),
}
cmd := exec.Command(s.ffprobePath,
"-v", "error",
"-show_format",
"-show_streams",
"-of", "json",
path,
)
output, err := cmd.Output()
if err != nil {
return meta, nil
}
var probe ffprobeOutput
if err := json.Unmarshal(output, &probe); err != nil {
return meta, nil
}
if probe.Format.FormatName != "" {
parts := strings.Split(probe.Format.FormatName, ",")
meta.Container = parts[0]
}
for _, stream := range probe.Streams {
switch stream.CodecType {
case "video":
vs := VideoStream{
Codec: stream.CodecName,
Profile: stream.Profile,
Width: stream.Width,
Height: stream.Height,
FPS: parseFPS(stream.RFrameRate),
PixelFormat: stream.PixFmt,
Bitrate: formatBitrate(stream.BitRate),
}
meta.VideoStreams = append(meta.VideoStreams, vs)
if stream.Width > 0 && stream.Height > 0 {
meta.Resolution = fmt.Sprintf("%dx%d", stream.Width, stream.Height)
}
case "audio":
as := AudioStream{
Codec: stream.CodecName,
SampleRate: stream.SampleRate,
Channels: stream.Channels,
ChannelLayout: stream.ChannelLayout,
SampleFormat: stream.SampleFmt,
Bitrate: formatBitrate(stream.BitRate),
Language: stream.Tags.Language,
}
meta.AudioStreams = append(meta.AudioStreams, as)
case "subtitle":
ss := SubtitleStream{
Codec: stream.CodecName,
Language: stream.Tags.Language,
Title: stream.Tags.Title,
}
meta.SubtitleStreams = append(meta.SubtitleStreams, ss)
}
}
return meta, nil
}
func formatBitrate(bitRate string) string {
if bitRate == "" {
return ""
}
br, err := strconv.ParseInt(bitRate, 10, 64)
if err != nil {
return ""
}
return fmt.Sprintf("%d", br/1000)
}
func parseFPS(rate string) string {
var num, den float64
if _, err := fmt.Sscanf(rate, "%f/%f", &num, &den); err != nil || den == 0 {
return ""
}
fps := num / den
if fps == math.Trunc(fps) {
return fmt.Sprintf("%.0f", fps)
}
return fmt.Sprintf("%.2f", fps)
}
type MediaMetadata struct {
Container string
Resolution string
FileSize int64
VideoStreams []VideoStream
AudioStreams []AudioStream
SubtitleStreams []SubtitleStream
}
type VideoStream struct {
Codec string
Profile string
Width int
Height int
FPS string
PixelFormat string
Bitrate string
}
type AudioStream struct {
Codec string
SampleRate string
Channels int
ChannelLayout string
SampleFormat string
Bitrate string
Language string
}
type SubtitleStream struct {
Language string
Title string
Codec string
}
type ffprobeOutput struct {
Format struct {
FormatName string `json:"format_name"`
BitRate string `json:"bit_rate"`
} `json:"format"`
Streams []ffprobeStream `json:"streams"`
}
type ffprobeStream struct {
Index int `json:"index"`
CodecName string `json:"codec_name"`
CodecType string `json:"codec_type"`
Profile string `json:"profile"`
Width int `json:"width"`
Height int `json:"height"`
RFrameRate string `json:"r_frame_rate"`
AvgFrameRate string `json:"avg_frame_rate"`
PixFmt string `json:"pix_fmt"`
SampleRate string `json:"sample_rate"`
Channels int `json:"channels"`
ChannelLayout string `json:"channel_layout"`
SampleFmt string `json:"sample_fmt"`
BitRate string `json:"bit_rate"`
Tags struct {
Language string `json:"language"`
Title string `json:"title"`
} `json:"tags"`
}
Ainternal/service/media_probe_test.go
@@ -0,0 +1,57 @@
package service
import (
"path/filepath"
"testing"
)
func TestParseFPS(t *testing.T) {
tests := []struct{ in, want string }{
{"30/1", "30"},
{"30000/1001", "29.97"},
{"0/0", ""},
{"", ""},
{"garbage", ""},
}
for _, tc := range tests {
if got := parseFPS(tc.in); got != tc.want {
t.Errorf("parseFPS(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
func TestFormatBitrate(t *testing.T) {
tests := []struct{ in, want string }{
{"128000", "128"},
{"", ""},
{"notanumber", ""},
}
for _, tc := range tests {
if got := formatBitrate(tc.in); got != tc.want {
t.Errorf("formatBitrate(%q) = %q, want %q", tc.in, got, tc.want)
}
}
}
func TestGetMetadataForSelectedFile(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "multi", "name = \"M\"\nduration = -1\n", nil)
makeTestVideoSize(t, filepath.Join(itemDir, "small.mp4"), "64x64")
makeTestVideoSize(t, filepath.Join(itemDir, "big.mp4"), "128x72")
m1, err := svc.GetMetadata("multi", "small.mp4")
if err != nil {
t.Fatalf("GetMetadata small: %v", err)
}
if m1.Resolution != "64x64" {
t.Errorf("small.mp4 resolution = %q, want 64x64", m1.Resolution)
}
m2, err := svc.GetMetadata("multi", "big.mp4")
if err != nil {
t.Fatalf("GetMetadata big: %v", err)
}
if m2.Resolution != "128x72" {
t.Errorf("big.mp4 resolution = %q, want 128x72", m2.Resolution)
}
}
Minternal/service/subscription.go
@@ -106,6 +106,17 @@ func (s *SubscriptionService) MarkRun(id int64, lastRunAt, nextRunAt time.Time,
return s.repo.MarkRun(id, lastRunAt, nextRunAt, status)
}
// MarkManualRun records a run started from the subscriptions page, leaving the
// schedule's next run time untouched.
func (s *SubscriptionService) MarkManualRun(id int64, lastRunAt time.Time, status string) error {
return s.repo.MarkManualRun(id, lastRunAt, status)
}
// SetLastStatus updates the recorded outcome of the current/last run.
func (s *SubscriptionService) SetLastStatus(id int64, status string) error {
return s.repo.SetLastStatus(id, status)
}
// Delete removes the subscription and the download-archive file "skip" mode
// keeps for it. Leaving the archive behind would make a later subscription that
// reuses the id silently skip entries it never downloaded. The library
Ainternal/service/subscription_run.go
@@ -0,0 +1,306 @@
package service
import (
"context"
"encoding/json"
"fmt"
"log"
"os"
"os/exec"
"path/filepath"
"strings"
"vidarchive/internal/models"
)
// refreshAndAddNew handles a metadata-mode run. The main pass used
// --skip-download, so tempDownloadDir holds only info.json files. Existing
// library items have their markers refreshed in place; entries with no existing
// match are genuinely new and are downloaded as full items in a second pass.
func (s *DownloadService) refreshAndAddNew(ctx context.Context, d *models.Download, preset *models.Preset, tempDownloadDir, ytdlpFlags string) error {
// The library dir is not created here: a refresh that matches everything
// writes nothing, and importDownloadedItems creates it when a second pass
// actually has an item to add.
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
return err
}
entries, err := os.ReadDir(tempDownloadDir)
if err != nil {
return err
}
var newURLs []string
for _, entry := range entries {
if err := ctx.Err(); err != nil {
return err
}
if !entry.IsDir() || !strings.HasPrefix(entry.Name(), "item-") {
continue
}
itemDir := filepath.Join(tempDownloadDir, entry.Name())
infoJSONPath := findInfoJSON(itemDir)
if infoJSONPath == "" {
continue
}
info := readInfoJSON(infoJSONPath)
if info.ID == "" {
continue
}
if existing, ok := s.librarySvc.FindByVideoID(baseLibraryDir, info.ID); ok {
if err := s.applyMetadata(existing, info, infoJSONPath); err != nil {
log.Printf("warning: failed to refresh metadata for %s: %v", itemDir, err)
}
continue
}
if info.WebpageURL != "" {
newURLs = append(newURLs, info.WebpageURL)
}
}
if len(newURLs) == 0 {
return nil
}
return s.downloadFresh(ctx, d, preset, newURLs, ytdlpFlags)
}
// downloadFresh fetches the given item URLs as full downloads (media + info.json)
// and imports them into the download's library directory. Metadata mode uses this
// to add entries that don't exist in the library yet.
func (s *DownloadService) downloadFresh(ctx context.Context, d *models.Download, preset *models.Preset, urls []string, ytdlpFlags string) error {
tempDir := s.tempNewDirFor(d.ID)
if err := os.MkdirAll(tempDir, 0755); err != nil {
return err
}
defer os.RemoveAll(tempDir)
args := s.presetSvc.BuildArgs(preset, d.FormatOverride, d.CustomFlags)
args, cleanup := s.appendCookies(args)
defer cleanup()
args = append(args, "--write-info-json")
args = append(args, "-P", tempDir)
args = append(args, "-o", "item-%(autonumber)05d/%(title)s.%(ext)s")
args = append(args, urls...)
runErr := s.runYTDLP(ctx, d, args)
// A cancelled second pass has nothing worth importing.
if ctx.Err() != nil {
return runErr
}
// Import whatever succeeded even if some entries errored.
if _, err := s.importDownloadedItems(ctx, d, tempDir, "", ytdlpFlags); err != nil {
log.Printf("warning: failed to import new metadata-mode items: %v", err)
}
return runErr
}
// mergeInfoJSON keeps fields from the existing sidecar that are absent from a
// metadata-only refresh. In particular, comments and heatmap data are expensive
// to reacquire and must not disappear just because the refresh preset does not
// request them.
func mergeInfoJSON(oldData, newData []byte) ([]byte, error) {
var oldObject, newObject map[string]json.RawMessage
if err := json.Unmarshal(newData, &newObject); err != nil {
return nil, err
}
if err := json.Unmarshal(oldData, &oldObject); err != nil {
return newData, nil
}
if newObject == nil {
return newData, nil
}
merged := make(map[string]json.RawMessage, len(oldObject)+len(newObject))
for key, value := range oldObject {
merged[key] = value
}
for key, value := range newObject {
merged[key] = value
}
for _, key := range []string{"comments", "heatmap"} {
oldValue, hadOldValue := oldObject[key]
newValue, hasNewValue := newObject[key]
if hadOldValue && (!hasNewValue || isEmptyJSONArray(newValue)) {
merged[key] = oldValue
}
}
return json.Marshal(merged)
}
func isEmptyJSONArray(value json.RawMessage) bool {
var values []json.RawMessage
if err := json.Unmarshal(value, &values); err != nil {
return false
}
return len(values) == 0
}
// restoreFileAtomically puts data back at path without exposing a partial file.
// It is used to roll back the marker if installing the staged info sidecar fails.
func restoreFileAtomically(path string, data []byte) error {
tmp, err := os.CreateTemp(filepath.Dir(path), ".vidarchive-restore-*.tmp")
if err != nil {
return err
}
tmpPath := tmp.Name()
defer os.Remove(tmpPath)
if err := tmp.Chmod(0644); err != nil {
tmp.Close()
return err
}
if _, err := tmp.Write(data); err != nil {
tmp.Close()
return err
}
if err := tmp.Close(); err != nil {
return err
}
if err := os.Rename(tmpPath, path); err != nil {
return err
}
return nil
}
// applyMetadata refreshes an existing item's marker and info sidecar from a
// fresh info.json without touching its media.
func (s *DownloadService) applyMetadata(existing string, info infoJSON, sourceInfoJSON string) error {
meta, err := s.librarySvc.readMetadata(existing)
if err != nil {
return fmt.Errorf("read metadata for %s: %w", existing, err)
}
if info.Title != "" {
meta.Name = info.Title
}
if info.Description != "" {
meta.Description = info.Description
}
if meta.SourceURL == "" && info.WebpageURL != "" {
meta.SourceURL = info.WebpageURL
}
meta.VideoID = info.ID
markerPath := filepath.Join(existing, itemMarkerName)
oldMarker, err := os.ReadFile(markerPath)
if err != nil {
return fmt.Errorf("read existing marker: %w", err)
}
// Stage the sidecar before changing the marker. The marker is committed first;
// if installing the sidecar then fails, restore the old marker so an ordinary
// I/O error cannot leave the two metadata files out of sync.
stagedInfo := ""
defer func() {
if stagedInfo != "" {
_ = os.Remove(stagedInfo)
}
}()
if sourceInfoJSON != "" {
data, err := os.ReadFile(sourceInfoJSON)
if err != nil {
return fmt.Errorf("read refreshed info JSON: %w", err)
}
if oldInfoJSON := findInfoJSON(existing); oldInfoJSON != "" {
oldData, err := os.ReadFile(oldInfoJSON)
if err != nil {
return fmt.Errorf("read existing info JSON: %w", err)
}
data, err = mergeInfoJSON(oldData, data)
if err != nil {
return fmt.Errorf("merge refreshed info JSON: %w", err)
}
}
tmp, err := os.CreateTemp(existing, ".info-json-*.tmp")
if err != nil {
return fmt.Errorf("create refreshed info JSON: %w", err)
}
stagedInfo = tmp.Name()
if err := tmp.Chmod(0644); err != nil {
tmp.Close()
return fmt.Errorf("set refreshed info JSON permissions: %w", err)
}
if _, err := tmp.Write(data); err != nil {
tmp.Close()
return fmt.Errorf("write refreshed info JSON: %w", err)
}
if err := tmp.Close(); err != nil {
return fmt.Errorf("close refreshed info JSON: %w", err)
}
}
if err := s.librarySvc.writeMetadata(existing, meta); err != nil {
return err
}
if stagedInfo != "" {
if err := os.Rename(stagedInfo, filepath.Join(existing, "info.json")); err != nil {
if restoreErr := restoreFileAtomically(markerPath, oldMarker); restoreErr != nil {
return fmt.Errorf("install refreshed info JSON: %v; restore marker: %w", err, restoreErr)
}
return fmt.Errorf("install refreshed info JSON: %w", err)
}
stagedInfo = ""
}
if rel, err := filepath.Rel(s.cfg.LibraryDir, existing); err == nil {
s.librarySvc.evictCachedScan(filepath.ToSlash(rel))
}
return nil
}
// pruneSubscription mirrors the source by deleting items in the subscription's
// directory that are no longer present upstream. It enumerates the current id
// set with a cheap flat-playlist listing; it never prunes when that enumeration
// fails or returns nothing, so a dead URL or network error can't wipe the dir.
func (s *DownloadService) pruneSubscription(ctx context.Context, d *models.Download, sub *models.Subscription) {
baseLibraryDir, err := s.resolveBaseLibraryDir(d)
if err != nil {
log.Printf("subscription %d prune skipped: %v", sub.ID, err)
return
}
keep, err := s.enumeratePlaylistIDs(ctx, sub.URL)
if err != nil {
log.Printf("subscription %d prune skipped: enumeration failed: %v", sub.ID, err)
return
}
if len(keep) == 0 {
log.Printf("subscription %d prune skipped: source returned no entries", sub.ID)
return
}
removed, err := s.librarySvc.PruneToIDSet(baseLibraryDir, keep)
if err != nil {
log.Printf("subscription %d prune error: %v", sub.ID, err)
return
}
if removed > 0 {
log.Printf("subscription %d pruned %d item(s) removed upstream", sub.ID, removed)
}
}
// enumeratePlaylistIDs lists the current video-id set for a URL without
// downloading, using yt-dlp --flat-playlist. Cookies are applied so private
// playlists enumerate correctly. Ids alone are sufficient to match items within
// a subscription's own directory (see FindByVideoID / PruneToIDSet).
func (s *DownloadService) enumeratePlaylistIDs(ctx context.Context, url string) (map[string]bool, error) {
args := []string{"--flat-playlist", "--no-warnings", "--print", "%(id)s"}
args, cleanup := s.appendCookies(args)
defer cleanup()
args = append(args, url)
out, err := exec.CommandContext(ctx, s.cfg.YTDLPPath, args...).Output()
if err != nil {
return nil, err
}
keep := make(map[string]bool)
for _, line := range strings.Split(string(out), "\n") {
id := strings.TrimSpace(line)
// yt-dlp prints "NA" for a missing field; never treat that as a real id.
if id == "" || id == "NA" {
continue
}
keep[id] = true
}
return keep, nil
}
Ainternal/service/subscription_run_test.go
@@ -0,0 +1,89 @@
package service
import (
"encoding/json"
"os"
"path/filepath"
"testing"
)
func TestApplyMetadataPreservesEngagementAndPermissions(t *testing.T) {
e := newExecEnv(t)
existing := writeItem(t, e.cfg.LibraryDir, "existing", `name = "Old"
video_id = "old"
`, map[string]string{
"video.mp4": "dummy",
"info.json": `{"id":"old","title":"Old","comments":[{"id":"c1","text":"archived"}],"heatmap":[{"start_time":0,"end_time":1,"value":1}],"filesize":123}`,
})
source := filepath.Join(e.scratch, "refreshed.info.json")
if err := os.WriteFile(source, []byte(`{"id":"new","title":"New","description":"updated","comments":[]}`), 0644); err != nil {
t.Fatal(err)
}
if err := e.svc.applyMetadata(existing, infoJSON{ID: "new", Title: "New", Description: "updated"}, source); err != nil {
t.Fatalf("applyMetadata: %v", err)
}
data, err := os.ReadFile(filepath.Join(existing, "info.json"))
if err != nil {
t.Fatal(err)
}
var got map[string]json.RawMessage
if err := json.Unmarshal(data, &got); err != nil {
t.Fatalf("parse installed info JSON: %v", err)
}
if _, ok := got["comments"]; !ok {
t.Fatal("refreshed sidecar lost archived comments")
}
if _, ok := got["heatmap"]; !ok {
t.Fatal("refreshed sidecar lost archived heatmap")
}
if string(got["title"]) != `"New"` {
t.Errorf("title = %s, want New", got["title"])
}
if mode := fileMode(t, filepath.Join(existing, "info.json")); mode != 0644 {
t.Errorf("info.json mode = %o, want 0644", mode)
}
}
func TestApplyMetadataRollsBackMarkerWhenSidecarInstallFails(t *testing.T) {
e := newExecEnv(t)
existing := writeItem(t, e.cfg.LibraryDir, "rollback", `name = "Old"
video_id = "old"
`, map[string]string{"video.mp4": "dummy"})
markerPath := filepath.Join(existing, itemMarkerName)
before, err := os.ReadFile(markerPath)
if err != nil {
t.Fatal(err)
}
// A directory at the destination makes the final rename fail after the
// marker has been written, exercising the rollback path.
infoPath := filepath.Join(existing, "info.json")
if err := os.Mkdir(infoPath, 0755); err != nil {
t.Fatal(err)
}
source := filepath.Join(e.scratch, "rollback.info.json")
if err := os.WriteFile(source, []byte(`{"id":"new","title":"New"}`), 0644); err != nil {
t.Fatal(err)
}
if err := e.svc.applyMetadata(existing, infoJSON{ID: "new", Title: "New"}, source); err == nil {
t.Fatal("expected sidecar installation to fail")
}
after, err := os.ReadFile(markerPath)
if err != nil {
t.Fatal(err)
}
if string(after) != string(before) {
t.Errorf("marker changed after failed sidecar install:\nbefore: %s\nafter: %s", before, after)
}
}
func fileMode(t *testing.T, path string) os.FileMode {
t.Helper()
info, err := os.Stat(path)
if err != nil {
t.Fatal(err)
}
return info.Mode().Perm()
}
Ainternal/service/subscription_status_test.go
@@ -0,0 +1,116 @@
package service
import (
"context"
"database/sql"
"testing"
"vidarchive/internal/models"
)
// newSubscriptionFor inserts a subscription owning dirName, the shape the
// scheduler creates downloads from.
func (e *execEnv) newSubscriptionFor(t *testing.T, dirName, refreshMode string) *models.Subscription {
t.Helper()
sub := &models.Subscription{
Name: "Channel",
URL: "https://example.com/channel",
Enabled: true,
RefreshMode: refreshMode,
ScheduleKind: "daily",
CronExpr: "0 3 * * *",
OutputDir: dirName,
}
if err := e.subRepo.Create(sub); err != nil {
t.Fatalf("create subscription: %v", err)
}
// The scheduler records "queued" when it enqueues the run; the worker is
// expected to move it on from there.
if err := e.subRepo.SetLastStatus(sub.ID, "queued"); err != nil {
t.Fatalf("seed last_status: %v", err)
}
return sub
}
func (e *execEnv) subStatus(t *testing.T, id int64) string {
t.Helper()
got, err := e.subRepo.GetByID(id)
if err != nil {
t.Fatalf("reload subscription %d: %v", id, err)
}
return got.LastStatus.String
}
// A finished subscription run must leave the subscription reporting the outcome.
// Without the write-back the row keeps the scheduler's "queued" forever, so the
// subscriptions page never shows that a run succeeded.
func TestSubscriptionRunRecordsCompleted(t *testing.T) {
e := newExecEnv(t)
sub := e.newSubscriptionFor(t, "channel", "overwrite")
e.fakeYTDLP(t, e.writeItem(t, "item-00001", "My Clip", "abc123"))
d := e.queue(t, &models.Download{SubscriptionID: sql.NullInt64{Int64: sub.ID, Valid: true}})
if _, err := e.svc.ExecuteDownload(context.Background(), d); err != nil {
t.Fatalf("ExecuteDownload: %v", err)
}
if got := e.subStatus(t, sub.ID); got != "completed" {
t.Errorf("subscription last_status = %q, want completed", got)
}
}
// A failed run must be visible on the subscription, not only on the download.
func TestSubscriptionRunRecordsError(t *testing.T) {
e := newExecEnv(t)
sub := e.newSubscriptionFor(t, "channel", "overwrite")
e.fakeYTDLP(t, `echo "boom" >&2; exit 1`)
d := e.queue(t, &models.Download{SubscriptionID: sql.NullInt64{Int64: sub.ID, Valid: true}})
if _, err := e.svc.ExecuteDownload(context.Background(), d); err == nil {
t.Fatal("expected the failing run to report an error")
}
if got := e.subStatus(t, sub.ID); got != "error" {
t.Errorf("subscription last_status = %q, want error", got)
}
}
// A user-initiated cancel is terminal, so the subscription reports it rather
// than staying on "downloading".
func TestSubscriptionRunRecordsCancelled(t *testing.T) {
e := newExecEnv(t)
sub := e.newSubscriptionFor(t, "channel", "overwrite")
e.fakeYTDLP(t, `touch "$SCRATCH/started"; sleep 300`)
d := e.queue(t, &models.Download{SubscriptionID: sql.NullInt64{Int64: sub.ID, Valid: true}})
done := make(chan struct{})
go func() {
defer close(done)
e.svc.ExecuteDownload(context.Background(), d)
}()
waitForFile(t, e.scratch+"/started")
e.svc.cancelDownload(d.ID)
<-done
if got := e.subStatus(t, sub.ID); got != "cancelled" {
t.Errorf("subscription last_status = %q, want cancelled", got)
}
}
// A plain download has no subscription to report to; the write-back must not
// try to touch one.
func TestPlainDownloadLeavesSubscriptionsAlone(t *testing.T) {
e := newExecEnv(t)
sub := e.newSubscriptionFor(t, "channel", "overwrite")
e.fakeYTDLP(t, e.writeItem(t, "item-00001", "My Clip", "abc123"))
d := e.queue(t, &models.Download{})
if _, err := e.svc.ExecuteDownload(context.Background(), d); err != nil {
t.Fatalf("ExecuteDownload: %v", err)
}
if got := e.subStatus(t, sub.ID); got != "queued" {
t.Errorf("subscription last_status = %q, want it untouched at queued", got)
}
}
Ainternal/service/subtitles.go
@@ -0,0 +1,159 @@
package service
import (
"encoding/json"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"vidarchive/internal/models"
"vidarchive/internal/util"
)
func (s *LibraryService) SubtitleDir(relPath string) string {
itemDir, err := s.resolveItemDir(relPath)
if err != nil {
return ""
}
return filepath.Join(itemDir, subtitlesDirName)
}
// GetSubtitlePath returns the .vtt path for a language. It errors when the item
// can't be resolved: joining onto an empty dir would yield a bare relative name
// that the caller would then serve relative to the process working directory.
func (s *LibraryService) GetSubtitlePath(relPath, lang string) (string, error) {
dir := s.SubtitleDir(relPath)
if dir == "" {
return "", fmt.Errorf("resolve subtitle dir for %q", relPath)
}
return filepath.Join(dir, lang+".vtt"), nil
}
func (s *LibraryService) GetSubtitles(relPath string) ([]models.SubtitleTrack, error) {
item, err := s.GetByRelPath(relPath)
if err != nil {
return nil, err
}
cacheDir := s.SubtitleDir(relPath)
entries, err := os.ReadDir(cacheDir)
if err == nil && len(entries) > 0 {
var tracks []models.SubtitleTrack
for _, entry := range entries {
if entry.IsDir() || filepath.Ext(entry.Name()) != ".vtt" {
continue
}
lang := strings.TrimSuffix(entry.Name(), ".vtt")
tracks = append(tracks, models.SubtitleTrack{
Lang: lang,
Label: lang,
Src: fmt.Sprintf("/media/item/%s/subtitles/%s", util.URLEncodePath(relPath), lang),
})
}
return tracks, nil
}
for _, mf := range item.MediaFiles {
if mf.IsAudio {
continue
}
streams, err := s.extractSubtitleInfo(mf.Filepath)
if err != nil || len(streams) == 0 {
continue
}
if err := os.MkdirAll(cacheDir, 0755); err != nil {
return nil, err
}
var tracks []models.SubtitleTrack
for _, stream := range streams {
lang := stream.Lang
if lang == "" {
lang = fmt.Sprintf("track%d", stream.Index)
}
outPath := filepath.Join(cacheDir, lang+".vtt")
if err := s.extractSubtitleToVTT(mf.Filepath, outPath, stream.Index); err != nil {
continue
}
tracks = append(tracks, models.SubtitleTrack{
Lang: lang,
Label: stream.Label,
Src: fmt.Sprintf("/media/item/%s/subtitles/%s", util.URLEncodePath(relPath), lang),
})
}
return tracks, nil
}
return nil, nil
}
type subtitleStream struct {
Index int
Lang string
Label string
}
func (s *LibraryService) extractSubtitleInfo(path string) ([]subtitleStream, error) {
cmd := exec.Command(s.ffprobePath,
"-v", "error",
"-show_streams",
"-select_streams", "s",
"-of", "json",
path,
)
output, err := cmd.Output()
if err != nil {
return nil, err
}
var probe struct {
Streams []struct {
Index int `json:"index"`
CodecName string `json:"codec_name"`
Tags struct {
Language string `json:"language"`
Title string `json:"title"`
} `json:"tags"`
} `json:"streams"`
}
if err := json.Unmarshal(output, &probe); err != nil {
return nil, err
}
var streams []subtitleStream
subIndex := 0
for _, stream := range probe.Streams {
switch stream.CodecName {
case "subrip", "ass", "ssa", "webvtt", "mov_text":
lang := stream.Tags.Language
if lang == "" {
lang = fmt.Sprintf("track%d", subIndex)
}
label := stream.Tags.Title
if label == "" {
label = strings.ToUpper(lang)
}
streams = append(streams, subtitleStream{
Index: subIndex,
Lang: lang,
Label: label,
})
subIndex++
}
}
return streams, nil
}
func (s *LibraryService) extractSubtitleToVTT(inputPath, outputPath string, streamIndex int) error {
cmd := exec.Command(s.ffmpegPath,
"-i", inputPath,
"-map", fmt.Sprintf("0:s:%d", streamIndex),
"-f", "webvtt",
outputPath,
"-y",
)
output, err := cmd.CombinedOutput()
if err != nil {
return fmt.Errorf("ffmpeg subtitle extraction failed: %w\nOutput: %s", err, string(output))
}
return nil
}
Ainternal/service/thumbnail.go
@@ -0,0 +1,263 @@
package service
import (
"encoding/json"
"fmt"
"log"
"os"
"os/exec"
"path/filepath"
"strings"
"sync"
"sync/atomic"
"vidarchive/internal/models"
"vidarchive/internal/util"
)
// maxConcurrentThumbnails caps how many ffmpeg extraction processes may run at
// once, so a freshly loaded library page (which fires one thumbnail request per
// visible item) cannot spawn an unbounded ffmpeg storm.
const maxConcurrentThumbnails = 3
// ThumbnailForFile returns the thumbnail for a specific media file within an
// item, extracting it on demand if needed. The bool is false when no thumbnail
// is available (file not found, audio-only, or extraction failed) so the caller
// can serve an icon. An empty/unmatched filename yields false — every thumbnail
// is keyed to a specific media file.
func (s *LibraryService) ThumbnailForFile(relPath, filename string) (string, bool) {
if filename == "" {
return "", false
}
item, err := s.GetByRelPath(relPath)
if err != nil {
return "", false
}
for i := range item.MediaFiles {
if item.MediaFiles[i].Filename != filename {
continue
}
mf := item.MediaFiles[i]
if path, ok := s.findExistingThumbnail(mf.Filepath); ok {
return path, true
}
return s.ensureThumbnailForFile(mf)
}
return "", false
}
// ensureThumbnailForFile returns an existing thumbnail for mf or extracts one,
// serializing concurrent extraction of the same file via a per-path mutex. It
// re-checks the disk under the lock so that whichever caller wins the race does
// the work and the rest reuse the result. A prior in-process failure short-
// circuits to avoid re-running ffmpeg on every request (see thumbFailed).
func (s *LibraryService) ensureThumbnailForFile(mf models.MediaFile) (string, bool) {
actual, _ := s.thumbLocks.LoadOrStore(mf.Filepath, &sync.Mutex{})
lock := actual.(*sync.Mutex)
lock.Lock()
defer lock.Unlock()
if path, ok := s.findExistingThumbnail(mf.Filepath); ok {
return path, true
}
// Trust a prior failure for this run rather than re-running ffmpeg every
// request; a restart clears thumbFailed and retries. Checked after the disk
// so a thumbnail that appears later (e.g. added manually) still wins.
if _, failed := s.thumbFailed.Load(mf.Filepath); failed {
return "", false
}
s.thumbSem <- struct{}{}
cur := atomic.AddInt32(&s.extractInFlight, 1)
for {
max := atomic.LoadInt32(&s.extractMaxConcurrent)
if cur <= max || atomic.CompareAndSwapInt32(&s.extractMaxConcurrent, max, cur) {
break
}
}
defer func() {
atomic.AddInt32(&s.extractInFlight, -1)
<-s.thumbSem
}()
atomic.AddInt32(&s.extractAttempts, 1)
path, err := s.extractThumbnail(mf)
if err != nil {
log.Printf("thumbnail extraction failed for %s: %v", mf.Filepath, err)
s.thumbFailed.Store(mf.Filepath, struct{}{})
return "", false
}
return path, true
}
func (s *LibraryService) findExistingThumbnail(path string) (string, bool) {
ext := filepath.Ext(path)
base := strings.TrimSuffix(path, ext) + ".thumbnail"
for _, candidate := range []string{base + ".webp", base + ".jpg", base + ".jpeg", base + ".png"} {
if info, err := os.Stat(candidate); err == nil && info.Size() > 0 {
return candidate, true
}
}
return "", false
}
// findImageAttachment returns the ordinal (0-based among attachment streams) of
// the best image attachment in a container — e.g. the cover.jpg/cover.webp that
// yt-dlp embeds into MKV with --embed-thumbnail — or -1 if there is none. Such
// covers are attachment streams, not attached_pic video streams, so they must be
// dumped with -dump_attachment rather than mapped like a normal stream.
func (s *LibraryService) findImageAttachment(path string) int {
cmd := exec.Command(s.ffprobePath, "-v", "error", "-show_streams", "-of", "json", path)
output, err := cmd.Output()
if err != nil {
return -1
}
var probe struct {
Streams []struct {
CodecType string `json:"codec_type"`
Tags struct {
Mimetype string `json:"mimetype"`
Filename string `json:"filename"`
} `json:"tags"`
} `json:"streams"`
}
if err := json.Unmarshal(output, &probe); err != nil {
return -1
}
best, bestScore, attachmentIdx := -1, 0, 0
for _, stream := range probe.Streams {
if stream.CodecType != "attachment" {
continue
}
idx := attachmentIdx
attachmentIdx++
if !strings.HasPrefix(strings.ToLower(stream.Tags.Mimetype), "image/") {
continue
}
score := 20
switch name := strings.ToLower(stream.Tags.Filename); {
case strings.Contains(name, "cover"):
score = 100
case strings.Contains(name, "thumbnail"), strings.Contains(name, "thumb"):
score = 80
case strings.Contains(name, "poster"):
score = 60
case strings.Contains(name, "art"):
score = 40
}
if score > bestScore {
best, bestScore = idx, score
}
}
return best
}
// thumbAttempt is one ffmpeg invocation that may produce a thumbnail at out.
type thumbAttempt struct {
label string
out string
args []string
}
func (s *LibraryService) extractThumbnail(mf models.MediaFile) (string, error) {
ext := filepath.Ext(mf.Filepath)
base := strings.TrimSuffix(mf.Filepath, ext) + ".thumbnail"
webpPath := base + ".webp"
jpgPath := base + ".jpg"
var attempts []thumbAttempt
// Collect every attempt's error so a genuine failure surfaces all of them
// rather than only the last fallback's stderr.
var attemptErrs []string
// Prefer an embedded image attachment (e.g. yt-dlp's cover.webp/cover.jpg in
// MKV): dump its raw bytes, then transcode to a canonical WebP.
if idx := s.findImageAttachment(mf.Filepath); idx >= 0 {
if rawPath, err := s.dumpAttachment(mf.Filepath, idx); err != nil {
attemptErrs = append(attemptErrs, fmt.Sprintf("attachment-dump: %v", err))
} else {
defer os.Remove(rawPath)
attempts = append(attempts, thumbAttempt{"attachment-webp", webpPath,
[]string{"-i", rawPath, "-c:v", "libwebp"}})
}
}
// Next try an embedded cover art video stream (attached_pic), e.g. mp3/mp4,
// falling back to jpeg if libwebp or webp encoding fails.
embedded := []string{"-i", mf.Filepath, "-map", "0:v", "-map", "-0:V", "-vframes", "1"}
attempts = append(attempts,
thumbAttempt{"embedded-webp", webpPath, append(embedded, "-c:v", "libwebp")},
thumbAttempt{"embedded-jpg", jpgPath, append(embedded, "-q:v", "2")},
)
if !mf.IsAudio {
seekTime := "00:00:01"
if mf.Duration > 0 {
seekTime = util.FormatClock(mf.Duration / 2)
}
frame := []string{"-ss", seekTime, "-i", mf.Filepath, "-vframes", "1"}
attempts = append(attempts,
thumbAttempt{"frame-webp", webpPath, append(frame, "-c:v", "libwebp")},
thumbAttempt{"frame-jpg", jpgPath, append(frame, "-q:v", "2")},
)
}
for _, a := range attempts {
path, err := s.tryWriteThumbnail(a.out, a.args)
if err == nil {
return path, nil
}
attemptErrs = append(attemptErrs, fmt.Sprintf("%s: %v", a.label, err))
}
if mf.IsAudio {
return "", fmt.Errorf("audio file has no thumbnail")
}
return "", fmt.Errorf("all thumbnail extraction attempts failed:\n%s", strings.Join(attemptErrs, "\n"))
}
// dumpAttachment extracts the raw bytes of attachment stream idx (e.g. the
// cover.jpg/cover.webp yt-dlp embeds into MKV with --embed-thumbnail) into a
// temp file and returns its path. The caller removes the file.
func (s *LibraryService) dumpAttachment(path string, idx int) (string, error) {
raw, err := os.CreateTemp("", "vidarchive-attachment-*")
if err != nil {
return "", err
}
rawPath := raw.Name()
raw.Close()
dumpArgs := []string{
fmt.Sprintf("-dump_attachment:t:%d", idx), rawPath,
"-i", path, "-y", "-t", "0", "-f", "null", "-",
}
if out, err := exec.Command(s.ffmpegPath, dumpArgs...).CombinedOutput(); err != nil {
os.Remove(rawPath)
return "", fmt.Errorf("%v\n%s", err, out)
}
return rawPath, nil
}
// tryWriteThumbnail runs ffmpeg with args to produce outputPath. The temp file
// keeps the final extension so ffmpeg can infer the output muxer (it cannot for
// a bare ".tmp" suffix), then is atomically renamed into place.
func (s *LibraryService) tryWriteThumbnail(outputPath string, args []string) (string, error) {
outExt := filepath.Ext(outputPath)
tmpPath := strings.TrimSuffix(outputPath, outExt) + ".tmp" + outExt
os.Remove(tmpPath)
cmd := exec.Command(s.ffmpegPath, append(args, tmpPath)...)
if output, err := cmd.CombinedOutput(); err != nil {
os.Remove(tmpPath)
return "", fmt.Errorf("ffmpeg failed: %v\n%s", err, string(output))
}
if info, err := os.Stat(tmpPath); err != nil || info.Size() == 0 {
os.Remove(tmpPath)
return "", fmt.Errorf("ffmpeg produced empty output")
}
if err := os.Rename(tmpPath, outputPath); err != nil {
os.Remove(tmpPath)
return "", err
}
return outputPath, nil
}
Ainternal/service/thumbnail_test.go
@@ -0,0 +1,287 @@
package service
import (
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"sync"
"sync/atomic"
"testing"
)
// makeVideoWithCoverAttachment writes an MKV at outPath with a 64x64 video and
// an embedded "cover.webp" image attachment of the given size — mirroring how
// yt-dlp embeds thumbnails into MKV (a true attachment stream, not an
// attached_pic video stream; note ffmpeg muxes an attached .png as a video
// stream, so webp/jpg must be used to get a real attachment).
func makeVideoWithCoverAttachment(t *testing.T, outPath, coverSize string) {
t.Helper()
scratch := t.TempDir()
base := filepath.Join(scratch, "base.mp4")
makeTestVideoSize(t, base, "64x64")
cover := filepath.Join(scratch, "cover.webp")
if out, err := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-f", "lavfi", "-i", "color=red:size="+coverSize+":duration=1",
"-frames:v", "1", cover, "-y").CombinedOutput(); err != nil {
t.Fatalf("make cover: %v\n%s", err, out)
}
if out, err := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-i", base, "-attach", cover, "-metadata:s:t:0", "mimetype=image/webp",
"-c", "copy", outPath, "-y").CombinedOutput(); err != nil {
t.Fatalf("attach cover: %v\n%s", err, out)
}
}
// probeImageSize returns the pixel dimensions of an image/video file.
func probeImageSize(t *testing.T, path string) (int, int) {
t.Helper()
out, err := exec.Command("ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_entries", "stream=width,height", "-of", "csv=p=0:s=x", path).Output()
if err != nil {
t.Fatalf("probe %s: %v", path, err)
}
var w, h int
if _, err := fmt.Sscanf(strings.TrimSpace(string(out)), "%dx%d", &w, &h); err != nil {
t.Fatalf("parse dimensions %q: %v", out, err)
}
return w, h
}
func TestThumbnailPrefersExistingGenerated(t *testing.T) {
svc, dir := newLibrary(t)
writeItem(t, dir, "item", "name = \"I\"\nduration = -1\n", map[string]string{
"video.mp4": "v",
"video.thumbnail.webp": "GENERATED",
})
path, ok := svc.ThumbnailForFile("item", "video.mp4")
if !ok {
t.Fatal("expected a thumbnail")
}
if !strings.HasSuffix(path, "video.thumbnail.webp") {
t.Errorf("expected generated thumbnail, got %q", path)
}
}
func TestThumbnailExtractsOnDemand(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
path, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("expected on-demand extraction to succeed")
}
if info, err := os.Stat(path); err != nil || info.Size() == 0 {
t.Fatalf("thumbnail file missing/empty: %v", err)
}
if !strings.Contains(filepath.Base(path), ".thumbnail.") {
t.Errorf("unexpected thumbnail name %q", path)
}
// No leftover temp files from the atomic-write path.
entries, _ := os.ReadDir(itemDir)
for _, e := range entries {
if strings.Contains(e.Name(), ".tmp") {
t.Errorf("leftover temp file %q", e.Name())
}
}
}
func TestThumbnailRetriesAfterDeletion(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
first, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("first extraction failed")
}
if err := os.Remove(first); err != nil {
t.Fatal(err)
}
// A failure/absence must not be cached permanently: re-request re-extracts.
second, ok := svc.ThumbnailForFile("vid", "vid.mp4")
if !ok {
t.Fatal("re-extraction after deletion failed (failure was cached)")
}
if info, err := os.Stat(second); err != nil || info.Size() == 0 {
t.Fatalf("re-extracted thumbnail missing/empty: %v", err)
}
}
func TestThumbnailConcurrentSingleExtraction(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "vid", "name = \"V\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "vid.mp4"))
var wg sync.WaitGroup
for i := 0; i < 8; i++ {
wg.Add(1)
go func() {
defer wg.Done()
if _, ok := svc.ThumbnailForFile("vid", "vid.mp4"); !ok {
t.Error("concurrent Thumbnail failed")
}
}()
}
wg.Wait()
// Exactly one generated thumbnail, no temp leftovers despite the race.
entries, _ := os.ReadDir(itemDir)
var thumbs int
for _, e := range entries {
if strings.Contains(e.Name(), ".thumbnail.") {
thumbs++
}
if strings.Contains(e.Name(), ".tmp") {
t.Errorf("leftover temp file %q", e.Name())
}
}
if thumbs != 1 {
t.Errorf("expected exactly 1 generated thumbnail, got %d", thumbs)
}
}
func TestThumbnailForFileIsPerFile(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "multi", "name = \"M\"\nduration = -1\n", nil)
makeTestVideo(t, filepath.Join(itemDir, "a.mp4"))
makeTestVideo(t, filepath.Join(itemDir, "b.mp4"))
pa, ok := svc.ThumbnailForFile("multi", "a.mp4")
if !ok {
t.Fatal("thumbnail for a.mp4 failed")
}
pb, ok := svc.ThumbnailForFile("multi", "b.mp4")
if !ok {
t.Fatal("thumbnail for b.mp4 failed")
}
if pa == pb {
t.Errorf("expected distinct per-file thumbnails, both = %q", pa)
}
if !strings.Contains(filepath.Base(pa), "a.thumbnail.") {
t.Errorf("a.mp4 thumbnail name = %q", filepath.Base(pa))
}
if !strings.Contains(filepath.Base(pb), "b.thumbnail.") {
t.Errorf("b.mp4 thumbnail name = %q", filepath.Base(pb))
}
// An unknown file yields no thumbnail (caller falls back to an icon).
if _, ok := svc.ThumbnailForFile("multi", "nope.mp4"); ok {
t.Error("unknown file should not produce a thumbnail")
}
// Every thumbnail is keyed to a specific file: an empty filename yields none.
if _, ok := svc.ThumbnailForFile("multi", ""); ok {
t.Error("empty filename should not produce a thumbnail")
}
}
func TestThumbnailConcurrencyBounded(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "many", "name = \"M\"\nduration = -1\n", nil)
const n = 8
for i := 0; i < n; i++ {
makeTestVideo(t, filepath.Join(itemDir, fmt.Sprintf("c%d.mp4", i)))
}
var wg sync.WaitGroup
for i := 0; i < n; i++ {
i := i
wg.Add(1)
go func() {
defer wg.Done()
svc.ThumbnailForFile("many", fmt.Sprintf("c%d.mp4", i))
}()
}
wg.Wait()
max := atomic.LoadInt32(&svc.extractMaxConcurrent)
if max > maxConcurrentThumbnails {
t.Errorf("peak concurrent extractions %d exceeded cap %d", max, maxConcurrentThumbnails)
}
if max < 1 {
t.Error("expected at least one extraction to run")
}
}
func TestThumbnailUsesEmbeddedAttachment(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "att", "name = \"A\"\nduration = -1\n", nil)
mkv := filepath.Join(itemDir, "v.mkv")
// 100x100 cover so it's distinguishable from a 64x64 video frame.
makeVideoWithCoverAttachment(t, mkv, "100x100")
if idx := svc.findImageAttachment(mkv); idx < 0 {
t.Fatal("findImageAttachment did not find the embedded cover")
}
path, ok := svc.ThumbnailForFile("att", "v.mkv")
if !ok {
t.Fatal("thumbnail extraction failed")
}
// The embedded cover (100x100) must be used in preference to a video frame
// (which would be 64x64) — this is the regression the refactor introduced.
if w, h := probeImageSize(t, path); w != 100 || h != 100 {
t.Errorf("thumbnail is %dx%d, expected 100x100 from the embedded cover (got a video frame instead)", w, h)
}
}
func TestThumbnailNegativeCacheSkipsReextraction(t *testing.T) {
svc, dir := newLibrary(t)
// A file ffmpeg cannot extract a thumbnail from: every attempt fails.
writeItem(t, dir, "bad", "name = \"B\"\nduration = -1\n", map[string]string{
"broken.mp4": "not actually a video",
})
if _, ok := svc.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("expected extraction to fail for a non-video file")
}
attempts1 := atomic.LoadInt32(&svc.extractAttempts)
if attempts1 == 0 {
t.Fatal("expected at least one extraction attempt")
}
// A second request is served from the in-process negative cache: no new
// ffmpeg attempt.
if _, ok := svc.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("expected the cached failure to persist")
}
if attempts2 := atomic.LoadInt32(&svc.extractAttempts); attempts2 != attempts1 {
t.Errorf("negative cache should prevent re-extraction; attempts %d -> %d", attempts1, attempts2)
}
// The cache is in-process only: a fresh service (≈ a restart) retries.
fresh := NewLibraryService(dir, "ffmpeg", "ffprobe")
if _, ok := fresh.ThumbnailForFile("bad", "broken.mp4"); ok {
t.Fatal("fresh service still fails (file is unextractable)")
}
if atomic.LoadInt32(&fresh.extractAttempts) == 0 {
t.Error("a fresh service should retry extraction, not inherit the negative cache")
}
}
func TestThumbnailAudioOnlyHasNone(t *testing.T) {
requireFFmpeg(t)
svc, dir := newLibrary(t)
itemDir := writeItem(t, dir, "aud", "name = \"A\"\nduration = -1\n", nil)
// A real audio file with no cover art.
cmd := exec.Command("ffmpeg", "-hide_banner", "-loglevel", "error",
"-f", "lavfi", "-i", "sine=frequency=440:duration=1",
filepath.Join(itemDir, "aud.mp3"), "-y")
if out, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("make audio: %v\n%s", err, out)
}
if path, ok := svc.ThumbnailForFile("aud", "aud.mp3"); ok {
t.Errorf("audio-only item should have no thumbnail, got %q", path)
}
}
Ainternal/service/ytdlp_flags.go
@@ -0,0 +1,60 @@
package service
import (
"fmt"
"strings"
)
// reservedFlags are yt-dlp options VidArchive always sets itself; user custom
// flags must not pass them (or a conflicting inverse). The value describes what
// the option controls, for the failure message.
var reservedFlags = map[string]string{
"-o": "the output template",
"--output": "the output template",
"-P": "the download path",
"--paths": "the download path",
"--cookies": "cookies (set these in Settings instead)",
"--no-cookies": "cookies (set these in Settings instead)",
"--newline": "progress output formatting (VidArchive sets this to stream logs)",
// These hand yt-dlp an arbitrary command or binary to run. VidArchive passes
// custom flags through verbatim, so allowing them would turn the preset form
// into remote command execution.
"--exec": "running external commands (not permitted)",
"--exec-before-download": "running external commands (not permitted)",
"--postprocessor-args": "post-processor arguments (not permitted)",
"--ppa": "post-processor arguments (not permitted)",
"--downloader": "selecting an external downloader (not permitted)",
"--external-downloader": "selecting an external downloader (not permitted)",
"--downloader-args": "external downloader arguments (not permitted)",
"--external-downloader-args": "external downloader arguments (not permitted)",
}
// reservedSubscriptionFlags are additionally reserved for subscription runs,
// where VidArchive drives info-json writing and the refresh mode.
var reservedSubscriptionFlags = map[string]string{
"--write-info-json": "info-json writing (needed to track item identity)",
"--no-write-info-json": "info-json writing (needed to track item identity)",
"--download-archive": "the download archive (managed by Skip mode)",
"--no-download-archive": "the download archive (managed by Skip mode)",
"--skip-download": "media downloading (managed by Metadata mode)",
"--no-skip-download": "media downloading (managed by Metadata mode)",
}
// checkReservedFlags rejects custom flags that clash with options VidArchive
// controls, naming the offender. It matches both "--flag" and "--flag=value".
func checkReservedFlags(customFlags string, isSubscription bool) error {
for _, tok := range strings.Fields(customFlags) {
// Both "--flag value" and "--flag=value" name the same option.
name, _, _ := strings.Cut(tok, "=")
desc, ok := reservedFlags[name]
if !ok && isSubscription {
desc, ok = reservedSubscriptionFlags[name]
}
if ok {
return fmt.Errorf("custom flag %q conflicts with VidArchive's handling of %s; remove it and try again", tok, desc)
}
}
return nil
}
Ainternal/service/ytdlp_flags_test.go
@@ -0,0 +1,45 @@
package service
import (
"testing"
)
func TestCheckReservedFlags(t *testing.T) {
cases := []struct {
name string
flags string
subscription bool
wantErr bool
}{
{"empty", "", false, false},
{"harmless", "--no-playlist --write-thumbnail", false, false},
{"output short", "-o foo.mp4", false, true},
{"output long", "--output foo.mp4", false, true},
{"output equals form", "--output=foo.mp4", false, true},
{"paths short", "-P /tmp", false, true},
{"cookies", "--cookies x.txt", false, true},
{"cookies inverse", "--no-cookies", false, true},
// Subscription-only reserved flags pass for normal downloads...
{"skip-download non-sub", "--skip-download", false, false},
{"write-info-json non-sub", "--write-info-json", false, false},
// ...but are rejected for subscription runs (and their inverses).
{"skip-download sub", "--skip-download", true, true},
{"no-skip-download sub", "--no-skip-download", true, true},
{"write-info-json sub", "--write-info-json", true, true},
{"no-write-info-json sub", "--no-write-info-json", true, true},
{"download-archive sub", "--download-archive a.txt", true, true},
// The "--flag=value" spelling names the same option as "--flag value",
// for the subscription-only table as well as the base one.
{"download-archive equals form sub", "--download-archive=a.txt", true, true},
// Base reserved flags still apply to subscriptions.
{"output sub", "-o x", true, true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
err := checkReservedFlags(tc.flags, tc.subscription)
if tc.wantErr != (err != nil) {
t.Errorf("checkReservedFlags(%q, %v) error = %v, wantErr %v", tc.flags, tc.subscription, err, tc.wantErr)
}
})
}
}
Ainternal/worker/scheduler_test.go
@@ -0,0 +1,263 @@
package worker
import (
"os"
"path/filepath"
"testing"
"time"
"vidarchive/internal/config"
"vidarchive/internal/database"
"vidarchive/internal/models"
"vidarchive/internal/repository"
"vidarchive/internal/service"
)
type schedEnv struct {
sched *Scheduler
pool *Pool
subSvc *service.SubscriptionService
subRepo *repository.SubscriptionRepository
repo *repository.DownloadRepository
}
// newTestScheduler builds a scheduler over a real database. The pool is created
// but never started, so a queued run stays in the buffer where a test can count
// it instead of being executed.
func newTestScheduler(t *testing.T) *schedEnv {
t.Helper()
root := t.TempDir()
cfg := &config.Config{
DBPath: ":memory:",
DataDir: root,
LibraryDir: filepath.Join(root, "library"),
TempDir: filepath.Join(root, "temp"),
YTDLPPath: "/bin/false",
FFmpegPath: "ffmpeg",
FFprobePath: "ffprobe",
}
for _, dir := range []string{cfg.LibraryDir, cfg.TempDir} {
if err := os.MkdirAll(dir, 0755); err != nil {
t.Fatal(err)
}
}
db, err := database.New(cfg)
if err != nil {
t.Fatalf("init db: %v", err)
}
t.Cleanup(func() { db.Close() })
downloadRepo := repository.NewDownloadRepository(db)
subRepo := repository.NewSubscriptionRepository(db)
subSvc := service.NewSubscriptionService(subRepo, cfg)
downloadSvc := service.NewDownloadService(
downloadRepo,
service.NewLibraryService(cfg.LibraryDir, cfg.FFmpegPath, cfg.FFprobePath),
service.NewPresetService(repository.NewPresetRepository(db)),
service.NewSettingsService(repository.NewSettingsRepository(db)),
subSvc,
cfg,
)
pool := New(downloadSvc, 1)
return &schedEnv{
sched: NewScheduler(subSvc, downloadSvc, pool, time.Minute),
pool: pool,
subSvc: subSvc,
subRepo: subRepo,
repo: downloadRepo,
}
}
func (e *schedEnv) newSub(t *testing.T, name string, sub *models.Subscription) *models.Subscription {
t.Helper()
sub.Name = name
if sub.URL == "" {
sub.URL = "https://example.com/" + name
}
if sub.RefreshMode == "" {
sub.RefreshMode = "overwrite"
}
if sub.ScheduleKind == "" {
sub.ScheduleKind = "daily"
}
if sub.CronExpr == "" {
sub.CronExpr = "0 3 * * *"
}
if sub.OutputDir == "" {
sub.OutputDir = name
}
if err := e.subRepo.Create(sub); err != nil {
t.Fatalf("create subscription %s: %v", name, err)
}
return sub
}
func (e *schedEnv) reload(t *testing.T, id int64) *models.Subscription {
t.Helper()
got, err := e.subRepo.GetByID(id)
if err != nil {
t.Fatalf("reload subscription %d: %v", id, err)
}
return got
}
// A restart must not fire every subscription that happens to have no next run
// recorded: backfill gives it a time without queueing anything.
func TestBackfillNextRunsDoesNotRun(t *testing.T) {
e := newTestScheduler(t)
sub := e.newSub(t, "channel", &models.Subscription{Enabled: true})
e.sched.backfillNextRuns()
got := e.reload(t, sub.ID)
if !got.NextRunAt.Valid {
t.Fatal("next_run_at was not backfilled")
}
if !got.NextRunAt.Time.After(time.Now()) {
t.Errorf("next_run_at = %v, want a future time", got.NextRunAt.Time)
}
if got.LastStatus.Valid {
t.Errorf("last_status = %q, want nothing: backfill must not run the subscription", got.LastStatus.String)
}
if len(e.pool.queue) != 0 {
t.Errorf("%d downloads queued by a backfill", len(e.pool.queue))
}
}
// A disabled subscription is not scheduled at all, so it needs no next run.
func TestBackfillSkipsDisabled(t *testing.T) {
e := newTestScheduler(t)
sub := e.newSub(t, "paused", &models.Subscription{Enabled: false})
e.sched.backfillNextRuns()
if got := e.reload(t, sub.ID); got.NextRunAt.Valid {
t.Errorf("next_run_at = %v, want none for a disabled subscription", got.NextRunAt.Time)
}
}
// An invalid cron expression must not stop the other subscriptions from being
// backfilled.
func TestBackfillContinuesPastInvalidSchedule(t *testing.T) {
e := newTestScheduler(t)
bad := e.newSub(t, "bad", &models.Subscription{
Enabled: true, ScheduleKind: "cron", CronExpr: "not a cron",
})
good := e.newSub(t, "good", &models.Subscription{Enabled: true})
e.sched.backfillNextRuns()
if got := e.reload(t, bad.ID); got.NextRunAt.Valid {
t.Error("a subscription with a broken schedule should get no next run")
}
if got := e.reload(t, good.ID); !got.NextRunAt.Valid {
t.Error("the valid subscription was skipped after the broken one")
}
}
// A due subscription is queued, recorded as such, and pushed out to its next
// scheduled time so the same tick doesn't pick it up again.
func TestCheckDueQueuesAndReschedules(t *testing.T) {
e := newTestScheduler(t)
sub := e.newSub(t, "channel", &models.Subscription{Enabled: true})
e.sched.checkDue()
got := e.reload(t, sub.ID)
if got.LastStatus.String != "queued" {
t.Errorf("last_status = %q, want queued", got.LastStatus.String)
}
if !got.NextRunAt.Valid || !got.NextRunAt.Time.After(time.Now()) {
t.Errorf("next_run_at = %v, want a future time", got.NextRunAt.Time)
}
if len(e.pool.queue) != 1 {
t.Fatalf("%d downloads queued, want 1", len(e.pool.queue))
}
d := <-e.pool.queue
if !d.SubscriptionID.Valid || d.SubscriptionID.Int64 != sub.ID {
t.Errorf("queued download is not tagged with subscription %d", sub.ID)
}
if d.URL != sub.URL {
t.Errorf("queued URL = %q, want %q", d.URL, sub.URL)
}
// The subscription is no longer due, so a second tick queues nothing.
e.sched.checkDue()
if len(e.pool.queue) != 0 {
t.Errorf("%d downloads queued on a tick where nothing was due", len(e.pool.queue))
}
}
// A run longer than the interval must not stack a second one. The subscription
// is marked "skipped" and pushed forward, keeping its previous run time.
func TestRunSkipsWhileAnEarlierRunIsActive(t *testing.T) {
e := newTestScheduler(t)
sub := e.newSub(t, "channel", &models.Subscription{Enabled: true})
lastRun := time.Now().Add(-2 * time.Hour).Truncate(time.Second)
if err := e.subRepo.MarkRun(sub.ID, lastRun, time.Now().Add(-time.Hour), "queued"); err != nil {
t.Fatal(err)
}
// An in-flight download for this subscription.
active := &models.Download{URL: sub.URL, Status: "downloading"}
active.SubscriptionID.Int64, active.SubscriptionID.Valid = sub.ID, true
if err := e.repo.Create(active); err != nil {
t.Fatal(err)
}
e.sched.checkDue()
got := e.reload(t, sub.ID)
if got.LastStatus.String != "skipped" {
t.Errorf("last_status = %q, want skipped", got.LastStatus.String)
}
if !got.LastRunAt.Time.Equal(lastRun) {
t.Errorf("last_run_at = %v, want the earlier run preserved at %v", got.LastRunAt.Time, lastRun)
}
if !got.NextRunAt.Time.After(time.Now()) {
t.Errorf("next_run_at = %v, want it moved into the future", got.NextRunAt.Time)
}
if len(e.pool.queue) != 0 {
t.Errorf("%d downloads queued while a run was already active", len(e.pool.queue))
}
}
// A broken schedule must not be retried every tick; the run is deferred instead.
func TestRunDefersBrokenSchedule(t *testing.T) {
e := newTestScheduler(t)
sub := e.newSub(t, "bad", &models.Subscription{
Enabled: true, ScheduleKind: "cron", CronExpr: "not a cron",
})
e.sched.checkDue()
got := e.reload(t, sub.ID)
if !got.NextRunAt.Valid || !got.NextRunAt.Time.After(time.Now().Add(23*time.Hour)) {
t.Errorf("next_run_at = %v, want it deferred about a day", got.NextRunAt.Time)
}
}
// Stop must return promptly and end the loop.
func TestSchedulerStop(t *testing.T) {
e := newTestScheduler(t)
e.sched.Start()
e.sched.Stop()
if err := e.sched.ctx.Err(); err == nil {
t.Error("scheduler context still live after Stop")
}
}
// An interval of zero or less would spin the ticker, so the constructor clamps it.
func TestNewSchedulerRejectsNonPositiveInterval(t *testing.T) {
e := newTestScheduler(t)
for _, interval := range []time.Duration{0, -time.Second} {
s := NewScheduler(e.subSvc, nil, e.pool, interval)
if s.interval != time.Minute {
t.Errorf("interval %v became %v, want 1m", interval, s.interval)
}
}
}