Feature Extraction
Transformers
Safetensors
English
move_temporal
action-recognition
kth
cnn
video
custom_code
Eval Results (legacy)
Instructions to use harpertoken/move with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harpertoken/move with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="harpertoken/move", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("harpertoken/move", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
language: en
|
| 4 |
+
library_name: transformers
|
| 5 |
+
tags:
|
| 6 |
+
- action-recognition
|
| 7 |
+
- kth
|
| 8 |
+
- cnn
|
| 9 |
+
- video
|
| 10 |
+
datasets:
|
| 11 |
+
- kth-action-recognition
|
| 12 |
+
metrics:
|
| 13 |
+
- accuracy
|
| 14 |
+
model-index:
|
| 15 |
+
- name: move
|
| 16 |
+
results:
|
| 17 |
+
- task:
|
| 18 |
+
type: video-classification
|
| 19 |
+
dataset:
|
| 20 |
+
name: KTH
|
| 21 |
+
type: kth-action-recognition
|
| 22 |
+
split: test
|
| 23 |
+
metrics:
|
| 24 |
+
- type: accuracy
|
| 25 |
+
value: 0.5046
|
| 26 |
+
---
|
| 27 |
+
|
| 28 |
+
# move
|
| 29 |
+
|
| 30 |
+
A small action classifier over KTH video: sixteen 64 by 64 grayscale frames in, one of six actions out. Each frame passes a two-conv CNN, the sixteen representations max-pool over time, and a linear head predicts walking, jogging, running, boxing, handwaving or handclapping. Trained for ten epochs with Adam at learning rate 1e-3, batch size 32, seed 0. Test accuracy 0.5046 on 216 clips from held-out subjects. CPU training.
|
| 31 |
+
|
| 32 |
+
No transformer here. The `PreTrainedModel` wrapper exists only so the weights serialize as `config.json` plus `model.safetensors` and load through `AutoModel` (including `trust_remote_code` via `auto_map`), the same arrangement as `wear`, `tone`, `pole` and `fuse`.
|
| 33 |
+
|
| 34 |
+
## Usage
|
| 35 |
+
|
| 36 |
+
```python
|
| 37 |
+
import numpy as np
|
| 38 |
+
from transformers import AutoModel
|
| 39 |
+
from modeling_move import MoveTemporal # registers the architecture
|
| 40 |
+
|
| 41 |
+
model = AutoModel.from_pretrained("harpertoken/move")
|
| 42 |
+
model.eval()
|
| 43 |
+
clip = np.zeros((16, 64, 64), dtype=np.float32)
|
| 44 |
+
print(model.predict_label(clip))
|
| 45 |
+
```
|
| 46 |
+
|
| 47 |
+
`predict_label` takes sixteen 64 by 64 float frames in range 0 to 1 and returns the action. Frames are uniformly sampled across the clip and converted to grayscale at 64 by 64, matching training exactly. Needs `torch` and `transformers`.
|
| 48 |
+
|
| 49 |
+
## Training
|
| 50 |
+
|
| 51 |
+
KTH actions via a public mirror, subjects 01 to 16 for training (383 clips) and 17 to 25 for testing (216 clips), so no person appears in both sets. One sequence is absent from the mirror, hence 383 rather than 384.
|
| 52 |
+
|
| 53 |
+
The gate for publishing this model was temporal gain over a single-frame baseline using the identical encoder and preprocessing, differing only in aggregation:
|
| 54 |
+
|
| 55 |
+
| Model | Seed 0 | Seed 1 |
|
| 56 |
+
|---|---|---|
|
| 57 |
+
| Single middle frame | 0.4074 | 0.3704 |
|
| 58 |
+
| Temporal mean-pool | 0.4398 | not run |
|
| 59 |
+
| Temporal max-pool | 0.5046 | 0.4907 |
|
| 60 |
+
|
| 61 |
+
Mean-pooling gained 0.032 on 216 clips against a standard error near 0.034, which is not significant, so it was not published. Max-pooling gains 0.097 and 0.120 across two seeds, replicating in direction and magnitude. The published weights are the seed-0 max-pool run.
|
| 62 |
+
|
| 63 |
+
Two things went wrong before it worked. A concat-style fusion of frame features sat at chance; the max over time learned, because the most distinctive pose matters more than the average. And per-sample disk reads dominated wall time until clips were preloaded into RAM.
|
| 64 |
+
|
| 65 |
+
The wrapper was checked for exact equivalence: identical predictions on test batches in eval mode, after catching that dropout made train-mode outputs differ.
|
| 66 |
+
|
| 67 |
+
No dataset is published alongside this model. KTH already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead.
|
| 68 |
+
|
| 69 |
+
## Limitations
|
| 70 |
+
|
| 71 |
+
0.5046 is modest, and the card states it plainly. Six balanced classes put chance at 0.167, so the model has learned real motion distinctions, but fine confusions (jogging versus running, handwaving versus handclapping) dominate the errors. Anything outside short grayscale clips of single actors is out of scope.
|